Echo-Memory: A Controlled Study of Memory in Action World Models
arXiv:2606. 09803v1 Announce Type: cross Abstract: We present \textbf{Echo-Memory}, a controlled study of memory mechanisms in action-conditioned world models.
R2M-Bench is a benchmark that evaluates revisit memory in interactive video world models by comparing a revisit pair to two control pairs from the same rollout: a gap‑matched non‑revisit pair and a short‑range pair. It introduces MemoryGain (MG) and Normalized Memory Ratio (NMR) to quantify the revisit advantage over generic temporal stability and normalize it by short‑to‑baseline dynamics. Across 300 instances and seven models, NMR correlates with human judgments and reduces the influence of slow‑motion artifacts, with DreamX‑World‑Memo achieving the highest NMR.
arXiv:2606. 09803v1 Announce Type: cross Abstract: We present \textbf{Echo-Memory}, a controlled study of memory mechanisms in action-conditioned world models.
LayerRecall is a memory router for autoregressive video diffusion that selectively retrieves and injects historical key/value states into specific layers of the model, based on the current context. It addresses the problem that existing memory mechanisms expose nonlocal history but do not guarantee effective use, by recognizing that different layers prefer current, recent, or distant context. The method, combined with Cross‑Horizon Prediction Matching, achieves state‑of‑the‑art long‑range consistency on MemoBench and MovieBench while maintaining local continuity and incurring negligible inference overhead.
arXiv:2608. 07408v1 Announce Type: cross Abstract: We study visual persistence in interactive video world models.
arXiv:2608.23565v1 Announce Type: new Abstract: An interactive world model must follow the user's actions, remember the places it has shown, and stream in real time. The tension is structural: contro...
An interactive world model must follow the user's actions, remember the places it has shown, and stream in real time. The tension is structural: control wants a short horizon, memory wants an unbounde...
arXiv:2607. 13305v1 Announce Type: cross Abstract: Benchmark accuracy in video large language models (LLMs) is often treated as evidence of visual understanding.
arXiv:2606. 14732v1 Announce Type: cross Abstract: Autoregressive video diffusion models enable streaming generation but often degrade over long rollouts: static scene layouts drift, while mechanisms that improve spatial stability tend to suppress motion, causing natural flows such as water, fire, or smoke to stagnate.
arXiv:2608.29904v1 Announce Type: new Abstract: Modern video generators routinely fail at physical dynamics: objects float, trajectories violate gravity, contacts vanish. Standard denoising and flow-...
arXiv:2608. 12939v1 Announce Type: new Abstract: Joint-embedding predictive architectures (JEPAs) learn world models that predict in a compact latent space rather than in pixels, reducing the pressure to model nuisance appearance.
The paper introduces Relax Forcing, a training‑free memory mechanism for autoregressive video diffusion that structures temporal context into Sink, Tail, and History frames. By selecting History frames with a relaxation criterion, the method reduces error accumulation and attention overhead while preserving motion dynamics. Experiments on VBench‑Long demonstrate that this structured memory improves long‑video generation quality over existing baselines.
arXiv:2606. 31672v1 Announce Type: cross Abstract: Despite rapid progress in interactive world models (IWMs), existing benchmarks evaluate action following only at trajectory level and ignore memory and interaction physics.
arXiv:2607. 22705v1 Announce Type: cross Abstract: Object-centric learning aims to represent scenes as objects whose properties can be reused in new combinations.