Echo-Memory: A Controlled Study of Memory in Action World Models
arXiv:2606. 09803v1 Announce Type: cross Abstract: We present \textbf{Echo-Memory}, a controlled study of memory mechanisms in action-conditioned world models.
arXiv:2606. 09803v1 Announce Type: cross Abstract: We present \textbf{Echo-Memory}, a controlled study of memory mechanisms in action-conditioned world models.
arXiv:2609.28466v1 Announce Type: new Abstract: Advances in generative models have improved video fidelity, enabling long-horizon generation, interactive world modeling, and evolving visual environme...
arXiv:2606. 31495v1 Announce Type: new Abstract: We study a single idea across two settings: that a prediction-error signal, computed by a small predictor over the latent space of a frozen encoder, can serve both as a gate on plasticity and as a substrate for metacognition.
arXiv:2607. 16292v1 Announce Type: cross Abstract: Brain-encoding foundation models predict fMRI responses to video, audio, and text well enough to win the Algonauts 2025 challenge.
LayerRecall is a memory router for autoregressive video diffusion that selectively retrieves and injects historical key/value states into specific layers of the model, based on the current context. It addresses the problem that existing memory mechanisms expose nonlocal history but do not guarantee effective use, by recognizing that different layers prefer current, recent, or distant context. The method, combined with Cross‑Horizon Prediction Matching, achieves state‑of‑the‑art long‑range consistency on MemoBench and MovieBench while maintaining local continuity and incurring negligible inference overhead.
R2M-Bench is a benchmark that evaluates revisit memory in interactive video world models by comparing a revisit pair to two control pairs from the same rollout: a gap‑matched non‑revisit pair and a short‑range pair. It introduces MemoryGain (MG) and Normalized Memory Ratio (NMR) to quantify the revisit advantage over generic temporal stability and normalize it by short‑to‑baseline dynamics. Across 300 instances and seven models, NMR correlates with human judgments and reduces the influence of slow‑motion artifacts, with DreamX‑World‑Memo achieving the highest NMR.
arXiv:2608.28609v1 Announce Type: cross Abstract: A personalized agent needs a user memory: a persistent model of who its user is. Today it is almost always text -- transcripts and captions retrieved...
arXiv:2607. 16292v4 Announce Type: replace-cross Abstract: Brain-encoding foundation models predict fMRI responses to video, audio and text well enough to win the Algonauts 2025 challenge.
The paper introduces Causal Memory Policy (CMP), a framework that identifies the utility of memories in memory‑augmented language models by intervening on retrieval rather than on storage. CMP reserves fixed context slots for memories sampled with known propensities and estimates utility using self‑normalized inverse propensity weighting, providing unbiased estimates and exact variance. Experiments show that CMP improves discrimination between required and non‑required memories and reveals that identified utility alone is insufficient for retention decisions across unseen queries.
arXiv:2609.40222v1 Announce Type: new Abstract: When a camera revisits a previously observed region, a video world model should reproduce what was there before. This requires both remembering past ob...
arXiv:2608. 07408v1 Announce Type: cross Abstract: We study visual persistence in interactive video world models.
arXiv:2609.13012v1 Announce Type: new Abstract: Vision-language models retain a substantial amount of pixel-decodable visual content in their visual key-value cache. We show, in our setting, that thi...