Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing autoregressive video diffusion pipelines carry forward observation history as conditioning context, which makes shared state difficult to maintain in multi-agent and multi-view settings.
arXiv:2603. 02697v2 Announce Type: replace-cross Abstract: This paper presents ShareVerse, a video generation framework enabling multi-agent shared world modeling, addressing the gap in existing works that lack support for unified shared world construction with multi-agent interaction.
By Jiayi Zhu, Jianing Zhang, Yiying Yang, Wei Cheng, Xiaoyun Yuan
arXiv:2608.31005v1 Announce Type: new
Abstract: Existing long-video agents acquire evidence through one uniform behavior, ignoring whether the required evidence is concentrated, requires broad occurr...
By Can Zhang, Baofeng Zhang, Xiaotian Han, Junyuan Shang, Yuchen Ding, Shuohuan Wang, Dianhai Yu, Ruirui Li
The paper introduces TetherMem, a training‑free, query‑aware memory router designed for streaming autoregressive video models that generate long videos in chunks. By separating subject and scene queries and modulating historical access with region‑ and age‑conditioned priors, TetherMem prevents the model from anchoring the scene to stale backgrounds and viewpoints, a problem termed memory‑anchored scene under‑progression. In blinded pairwise evaluations, TetherMem outperforms eight baseline methods in overall quality and scene progression, and on full 30‑second videos it maintains background, viewpoint, and scene changes while preserving subject identity and temporal continuity.
By Chen Li, Peng Zhang, Hanyu Zhou, Jialong Zuo, Fei Wang, Daiguo Zhou, Nong Sang, Changxin Gao
arXiv:2610.02153v1 Announce Type: new
Abstract: Long-horizon autoregressive video generation is limited by a finite context window. When an object or scene falls out of context, its fine-grained visu...
By Yiwen Zhang, Haocheng Xi, Michael Tian-Yue Liu, Alexei A. Efros, Hadar Averbuch-Elor, Qianqian Wang, Haiwen Feng
arXiv:2606. 07649v1 Announce Type: cross Abstract: Long-form video generation requires systematic narrative planning and visual consistency that current short-clip methods cannot provide.
By Lingxuan Huang, Sizhe He, Hengji Zhou, Liqiang Nie, Lianghao Xia, Chao Huang
World in World introduces a training‑free inference interface that lets users control autoregressive video world models from new viewpoints. By converting diverse control signals—source‑video observations, target‑view projections, geometry renderings, and retrieved states—into camera‑ and time‑labelled visual tokens, the system uses a frozen causal video model’s self‑attention to maintain synchronization, complete unseen regions, and recover past appearances. The method supports camera‑controlled rerendering, long‑horizon revisiting, and human‑motion transfer while preserving perceptual quality, temporal consistency, and camera‑following accuracy.
MVAgent is a multi‑agent pipeline for multi‑shot video generation that ensures consistent character appearance, stable spatial layout, and continuous character state across shots. The system uses typed conditioning inputs: a Spatial Grounding agent samples camera views, an Observer records shot endings into a continuity memory, a Transition agent builds action and spatial references for subsequent shots, and an Orchestrator composes these inputs into generator requests. Trained with agentic reinforcement learning (Trunk‑GDPO) while keeping the generator and judges frozen, MVAgent achieves the highest cross‑shot consistency and narrative‑planning quality on ViMax‑Bench and is preferred over the strongest agentic baseline in human evaluation.
By Xiangyu Kong, Wenjie Zhou, Fengping Tian, Lihua Fang, Haoqin Sun, Chenyang Lyu, Longyue Wang, Weihua Luo
arXiv:2606. 02753v1 Announce Type: cross Abstract: Video world models are a foundational generative technology for embodied AI and the Metaverse, yet existing approaches are inherently limited to a single agent observing from a single perspective.
By Teng Hu, Mingchun Lu, Yating Wang, Jiangning Zhang, Jinkun Hao, Ye Pan, Ran Yi, Lizhuang Ma, Dacheng Tao
arXiv:2606. 16353v1 Announce Type: cross Abstract: Streaming video understanding models must answer queries at any moment during an ongoing stream, using only what they have observed so far and under fixed memory and computation budgets.
By Haonan Ge, Yiwei Wang, Hang Wu, Yujun Cai
arXiv:2608. 08612v1 Announce Type: cross Abstract: Recently, retrieval-augmented and memory-augmented methods have emerged as two promising paradigms for long-video question answering.
By Caijun Yan, Yang Zhou, Meixing Shi, Haoran Sun, Yichen Li, Yuxiang Cai, Yankai Jiang
arXiv:2608.05592v2 Announce Type: replace
Abstract: Multimodal Large Language Models (MLLMs) have made strong progress in video understanding, yet long videos remain difficult: the visual token budge...
By Ziling Huang, Shin'ichi Satoh