At ultra-low bitrates, high-fidelity reconstruction requires sampling plausible videos from the posterior rather than regressing to oversmoothed conditional means. GVCC (Generative Video Codebook Codec) turns a pretrained rectified-flow video generative model into a zero-shot decoder: the deterministic flow sampler is converted into an equivalent marginal-preserving stochastic process, and the transmitted bitstream encodes the per-step stochastic innovations that steer generation.
- GVCC works in three practical modes: Text-to-Video (T2V), autoregressive Image-to-Video (I2V) with tail latent correction, and First-Last-Frame-to-Video (FLF2V) with boundary-sharing GOP chaining. (NeurIPS 2026)
- GVCCTurbo is a BPP-driven scheduler that separates expensive prior refreshes from codebook corrections, turning bitrate into a schedule input and cutting prior evaluations from 20 to 9 (~44% decoding-time reduction).