Addressable Memory for Video World Models
arXiv:2608. 07408v1 Announce Type: cross Abstract: We study visual persistence in interactive video world models.
The paper "Can 4D Foundation Models Remember?" introduces PersistBench, a dataset and metric suite that uses 360° videos to evaluate visual memory in 4D foundation models. It focuses on three aspects—object permanence, motion continuity, and appearance preservation—to assess how well models remember objects after they leave the field of view. Experiments show that current models only maintain short‑term consistency, revealing a significant gap between perception and robust memory.
arXiv:2608. 07408v1 Announce Type: cross Abstract: We study visual persistence in interactive video world models.
arXiv:2608. 13492v1 Announce Type: new Abstract: This report presents an improved version of AlayaWorld.
arXiv:2610.02160v1 Announce Type: new Abstract: Precise control over camera and object motion is essential for professional video production. Existing methods control objects only coarsely, through i...
LayerRecall is a memory router for autoregressive video diffusion that selectively retrieves and injects historical key/value states into specific layers of the model, based on the current context. It addresses the problem that existing memory mechanisms expose nonlocal history but do not guarantee effective use, by recognizing that different layers prefer current, recent, or distant context. The method, combined with Cross‑Horizon Prediction Matching, achieves state‑of‑the‑art long‑range consistency on MemoBench and MovieBench while maintaining local continuity and incurring negligible inference overhead.
Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints. We present WorldCrafter, a video world m...
arXiv:2603. 03482v2 Announce Type: replace-cross Abstract: Interactive world models continually generate video by responding to a user's actions, enabling open-ended generation capabilities.
WorldCrafter is a video world model that introduces a camera‑queryable implicit 3D‑aware memory to improve long‑horizon consistency and viewpoint control. The model compresses multi‑view evidence into a limited token budget shaped by the requested viewpoint, integrating historical observations via a memory encoder and pose‑conditioned readout before denoising. Experiments on static and dynamic scenes demonstrate significant gains in consistency and camera‑control accuracy while maintaining visual quality during minute‑scale exploration.
arXiv:2608. 11017v1 Announce Type: cross Abstract: Long-horizon egocentric video is a rich substrate for wearable AI assistants, but object-centric questions such as where an item was moved, when it last changed state, or why it was relocated remain difficult because caption- and transcript-based memories rarely preserve persistent object identity or structured spatial change.
Long-horizon egocentric video is a rich substrate for wearable AI assistants, but object-centric questions such as where an item was moved, when it last changed state, or why it was relocated remain difficult because caption- and transcript-based memories rarely preserve persistent object identity or structured spatial change. Existing long-video QA methods mainly emphasize temporal grounding and clip retrieval, while prior 3D scene-graph methods typically assume stronger geometry than free-motion wearable RGB video provides, including point clouds, RGB-D input, posed views, sparse reconstruction, or reconstructed scenes.
arXiv:2605.25333v3 Announce Type: replace Abstract: Video world models should maintain evolving states when evidence is unobserved, yet current generators often freeze hidden states upon interruption...
arXiv:2609.35734v2 Announce Type: replace Abstract: Novel view synthesis from sparse images must reconcile faithful reconstruction of observed regions with plausible completion of unseen content, whi...
Real-world spatial intelligence requires agents to understand scenes from continuous video streams, where objects move, persist, disappear, and reappear over time. While recent spatial foundation models have enabled generalizable feed-forward 3D reconstruction, most streaming methods remain geometry-centric and lack temporally consistent object-level understanding.