arXiv Computer Vision

ConsistWorld: Evidence Routing for Consistent Multi-Agent World Models

Hugging Face Trending Papers
Jul 23

Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers

Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing autoregressive video diffusion pipelines carry forward observation history as conditioning context, which makes shared state difficult to maintain in multi-agent and multi-view settings.

arXiv Computer Vision
Aug 28

Tether the Subject, Release the Scene: Query-Aware Memory Routing for Long-Horizon Autoregressive Video Generation

The paper introduces TetherMem, a training‑free, query‑aware memory router designed for streaming autoregressive video models that generate long videos in chunks. By separating subject and scene queries and modulating historical access with region‑ and age‑conditioned priors, TetherMem prevents the model from anchoring the scene to stale backgrounds and viewpoints, a problem termed memory‑anchored scene under‑progression. In blinded pairwise evaluations, TetherMem outperforms eight baseline methods in overall quality and scene progression, and on full 30‑second videos it maintains background, viewpoint, and scene changes while preserving subject identity and temporal continuity.

By Chen Li, Peng Zhang, Hanyu Zhou, Jialong Zuo, Fei Wang, Daiguo Zhou, Nong Sang, Changxin Gao
Hugging Face Trending Papers
Sep 10

World in World: Explore the World with World Models

World in World introduces a training‑free inference interface that lets users control autoregressive video world models from new viewpoints. By converting diverse control signals—source‑video observations, target‑view projections, geometry renderings, and retrieved states—into camera‑ and time‑labelled visual tokens, the system uses a frozen causal video model’s self‑attention to maintain synchronization, complete unseen regions, and recover past appearances. The method supports camera‑controlled rerendering, long‑horizon revisiting, and human‑motion transfer while preserving perceptual quality, temporal consistency, and camera‑following accuracy.

arXiv Computer Vision
6d ago

MVAgent: Multi-Agent Video Generation via Consistent Condition Construction and Shot-Level Policy Optimization

MVAgent is a multi‑agent pipeline for multi‑shot video generation that ensures consistent character appearance, stable spatial layout, and continuous character state across shots. The system uses typed conditioning inputs: a Spatial Grounding agent samples camera views, an Observer records shot endings into a continuity memory, a Transition agent builds action and spatial references for subsequent shots, and an Orchestrator composes these inputs into generator requests. Trained with agentic reinforcement learning (Trunk‑GDPO) while keeping the generator and judges frozen, MVAgent achieves the highest cross‑shot consistency and narrative‑planning quality on ViMax‑Bench and is preferred over the strongest agentic baseline in human evaluation.

By Xiangyu Kong, Wenjie Zhou, Fengping Tian, Lihua Fang, Haoqin Sun, Chenyang Lyu, Longyue Wang, Weihua Luo