arXiv:2609.15863v1 Announce Type: new
Abstract: Video diffusion models are stochastic and hard to control: precise content often requires repeated sampling without guaranteed success, and long-horizo...
By Xiaofeng Mao, Peijia Lin, Shaohao Rui, Yibo Zhang, Haibin Wan, Weijie Ma
arXiv:2606. 09056v1 Announce Type: cross Abstract: Video generative models have become increasingly powerful, but long-range consistency remains challenging to achieve because even a few dozen frames require impractically long transformer sequence lengths.
By Ishaan Preetam Chandratreya, David Charatan, Basile Van Hoorick, Sergey Zakharov, Vitor Guizilini, Phillip Isola, Vincent Sitzmann
PixelUMM is an encoder‑free model that unifies image and video understanding and generation directly in pixel space. It represents images as spatial patches and videos as spatiotemporal tubelets, feeding both through single‑layer linear projections into a shared multimodal backbone. The Mixture‑of‑Transformers architecture blends shared attention with task‑specific parameters, enabling autoregressive text prediction, pixel‑space flow matching, and clean‑pixel video generation, and experiments show competitive performance across tasks while providing design insights for future pixel‑space multimodal models.
By Cong Wei, Xuanchi Ren, Bryan Chu, Weiming Ren, Huan Ling, Jiahui Huang, Laura Leal-Taix\'e, Sanja Fidler, Wenhu Chen, Zian Wang, Jay Zhangjie Wu
arXiv:2508.15774v2 Announce Type: replace
Abstract: Visual diffusion models achieve remarkable progress, yet they are typically trained at limited resolutions due to the lack of high-resolution data...
By Gordon Chen, Haonan Qiu, Ning Yu, Ziqi Huang, Paul Debevec, Ziwei Liu
arXiv:2607.14935v2 Announce Type: replace
Abstract: Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world application...
By Xinhao Li, Yuhan Zhu, Xiangyu Zeng, Yuhao Dong, Haoning Wu, Zhiqiu Zhang, Yuandong Yang, Changlian Ma, Qingyu Zhang, Yansong Shi, Xinyu Chen, Haoran Chen, Zizheng Huang, Jun Zhang, Kun Ouyang, Lin Sui, Ziang Yan, Yicheng Xu, Chenting Wang, Yinan He, Hongjie Zhang, Yi Wang, Yu Qiao, Yali Wang, Ziwei Liu, Kai Chen, Limin Wang
arXiv:2607. 18367v1 Announce Type: new Abstract: Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly.
By AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Mingliang Zhai, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li, Zhengyuan Lin, Zhixiang Wang, Zian Meng, Zihui Gao