arXiv AI By Jiancheng Huang, Gengwei Zhang, Zequn Jie, Siyu Jiao, Yinlong Qian, Ling Chen, Yunchao Wei, Lin Ma

M4V: Multimodal Mamba for Efficient Text-to-Video Generation

Read the original on arXiv AI →

arXiv:2506. 10915v2 Announce Type: replace-cross Abstract: Text-to-video generation has significantly enriched content creation and holds the potential to evolve into powerful world simulators.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jun 9

MilliVid: Hierarchical Latents for Long-Range Consistency in Video Generation

arXiv:2606. 09056v1 Announce Type: cross Abstract: Video generative models have become increasingly powerful, but long-range consistency remains challenging to achieve because even a few dozen frames require impractically long transformer sequence lengths.

By Ishaan Preetam Chandratreya, David Charatan, Basile Van Hoorick, Sergey Zakharov, Vitor Guizilini, Phillip Isola, Vincent Sitzmann
arXiv Computer Vision
3d ago

PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation

PixelUMM is an encoder‑free model that unifies image and video understanding and generation directly in pixel space. It represents images as spatial patches and videos as spatiotemporal tubelets, feeding both through single‑layer linear projections into a shared multimodal backbone. The Mixture‑of‑Transformers architecture blends shared attention with task‑specific parameters, enabling autoregressive text prediction, pixel‑space flow matching, and clean‑pixel video generation, and experiments show competitive performance across tasks while providing design insights for future pixel‑space multimodal models.

By Cong Wei, Xuanchi Ren, Bryan Chu, Weiming Ren, Huan Ling, Jiahui Huang, Laura Leal-Taix\'e, Sanja Fidler, Wenhu Chen, Zian Wang, Jay Zhangjie Wu
arXiv Computer Vision
Aug 25

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

arXiv:2607.14935v2 Announce Type: replace Abstract: Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world application...

By Xinhao Li, Yuhan Zhu, Xiangyu Zeng, Yuhao Dong, Haoning Wu, Zhiqiu Zhang, Yuandong Yang, Changlian Ma, Qingyu Zhang, Yansong Shi, Xinyu Chen, Haoran Chen, Zizheng Huang, Jun Zhang, Kun Ouyang, Lin Sui, Ziang Yan, Yicheng Xu, Chenting Wang, Yinan He, Hongjie Zhang, Yi Wang, Yu Qiao, Yali Wang, Ziwei Liu, Kai Chen, Limin Wang
arXiv AI
Jul 22

AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

arXiv:2607. 18367v1 Announce Type: new Abstract: Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly.

By AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Mingliang Zhai, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li, Zhengyuan Lin, Zhixiang Wang, Zian Meng, Zihui Gao