ConsistWorld: Evidence Routing for Consistent Multi-Agent World Models
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing autoregressive video diffusion pipelines carry forward observation history as conditioning context, which makes shared state difficult to maintain in multi-agent and multi-view settings.
arXiv:2603. 02697v2 Announce Type: replace-cross Abstract: This paper presents ShareVerse, a video generation framework enabling multi-agent shared world modeling, addressing the gap in existing works that lack support for unified shared world construction with multi-agent interaction.
arXiv:2608.31005v1 Announce Type: new Abstract: Existing long-video agents acquire evidence through one uniform behavior, ignoring whether the required evidence is concentrated, requires broad occurr...
The paper introduces TetherMem, a training‑free, query‑aware memory router designed for streaming autoregressive video models that generate long videos in chunks. By separating subject and scene queries and modulating historical access with region‑ and age‑conditioned priors, TetherMem prevents the model from anchoring the scene to stale backgrounds and viewpoints, a problem termed memory‑anchored scene under‑progression. In blinded pairwise evaluations, TetherMem outperforms eight baseline methods in overall quality and scene progression, and on full 30‑second videos it maintains background, viewpoint, and scene changes while preserving subject identity and temporal continuity.
arXiv:2610.02153v1 Announce Type: new Abstract: Long-horizon autoregressive video generation is limited by a finite context window. When an object or scene falls out of context, its fine-grained visu...
arXiv:2606. 07649v1 Announce Type: cross Abstract: Long-form video generation requires systematic narrative planning and visual consistency that current short-clip methods cannot provide.