arXiv:2607.14088v2 Announce Type: replace
Abstract: Video generation models typically rely on 3D-VAEs trained for pixel-level reconstruction, whose latent spaces may underrepresent semantic structure...
By Zhihao Xie, Junfeng Wu, Xinting Hu, Junchao Huang, Li Jiang
arXiv:2610.00686v1 Announce Type: new
Abstract: Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffusion models. The choice of scene...
By Mikhail Dereviannykh, Vikram Voleti, Simon Donne, Mallikarjun Byrasandra Ramalinga Reddy, Shimon Vainer, Mark Boss
The paper introduces GVCC, a zero‑shot video compression framework that uses a pretrained generative video model as the decoder. GVCC transforms deterministic rectified‑flow samplers into stochastic processes, enabling the transmission of compressed information through per‑step stochastic innovations. The authors evaluate three GVCC variants—Text‑to‑Video, Image‑to‑Video, and First‑Last‑Frame‑to‑Video—on the UVG dataset, reporting perceptual, fidelity, and temporal metrics without claiming global rate‑distortion gains.
By Ziyue Zeng, Xun Su, Haoyuan Liu, Bingyu Lu, Yui Tatsumi, Hiroshi Watanabe
arXiv:2608.24293v1 Announce Type: new
Abstract: Latent diffusion models have emerged as a dominant framework for high-fidelity image and video synthesis, operating in compact latent spaces with varia...
By Yeonkyeong Lee, Hyunsung Go, Jongmin Kim, Sewoong Lim, Donghoon Lee
4D generation synthesizes dynamic 3D scenes from conditions such as text or images. Existing methods either reconstruct generated RGB videos with a separate 4D model or adapt a particular video generator to predict geometry directly.
arXiv:2606. 09056v1 Announce Type: cross Abstract: Video generative models have become increasingly powerful, but long-range consistency remains challenging to achieve because even a few dozen frames require impractically long transformer sequence lengths.
By Ishaan Preetam Chandratreya, David Charatan, Basile Van Hoorick, Sergey Zakharov, Vitor Guizilini, Phillip Isola, Vincent Sitzmann
arXiv:2512.21004v2 Announce Type: replace
Abstract: Recent advances in pretraining general foundation models have significantly improved performance across diverse downstream tasks. While autoregress...
By Jinghan Li, Yang Jin, Hao Jiang, Yadong Mu, Yang Song, Kun Xu
arXiv:2606. 13289v1 Announce Type: cross Abstract: Holistic visual tokenizers are fundamental to unified multimodal models (UMMs) as they map diverse visual inputs into a unified representation space.
By Guozhen Zhang, Xuerui Qiu, Yutao Cui, Tianhui Song, Changlin Li, Junzhe Li, Tao Huang, Xiao Zhang, Yang Li, Jianbing Wu, Miles Yang, Zhao Zhong, Liefeng Bo, Limin Wang
arXiv:2608. 08676v1 Announce Type: cross Abstract: Semantic vision encoders have become a central visual interface for multimodal understanding and semantic conditioning in image generation.
By Jinbo Yan, Limeng Qiao, Jie Qin, Junyan He, Feize Wu, Guanglu Wan
arXiv:2610.01942v1 Announce Type: new
Abstract: Predicting the future evolution of a scene is a fundamental capability for world modeling. Recent work has shown that operating in the feature space of...
By Efstathios Karypidis, Spyros Gidaris, Nikos Komodakis
arXiv:2609.40037v1 Announce Type: new
Abstract: Few-step autoregressive video generation enables efficient streaming synthesis, but errors introduced in early temporal blocks are reused as context an...
By Fangyu Lin, Xingtong Ge, Lunjie Zhu, Yi Zhang, Zhening Liu, Tianhang Wang, Mengfei Li, Yumeng Zhang, Guanglu Song, Yu Liu, Jun Zhang
TT-VidT is a video pretraining method that decouples the temporal axis by combining a per‑frame ViT-B/16 spatial encoder with a compact Temporal Transfer Layer trained via Diff Compression. The authors conduct a systematic 24‑configuration study to isolate architecture, objective, and decoder effects, showing that the full TT-VidT design yields the strongest motion‑sensitive representations. In downstream fine‑tuning, TT‑VidT outperforms state‑of‑the‑art baselines on Jester, Something‑Something V2, ARID, and Diving48 while using significantly fewer encoder FLOPs.
By Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai