Addressable Memory for Video World Models
arXiv:2608. 07408v1 Announce Type: cross Abstract: We study visual persistence in interactive video world models.
arXiv:2608. 07408v1 Announce Type: cross Abstract: We study visual persistence in interactive video world models.
arXiv:2608.22725v1 Announce Type: new Abstract: Short-drama generation has grown into a large, industrialized pipeline, and as it scales from isolated shots to the episode level, visual continuity ha...
arXiv:2606. 09803v1 Announce Type: cross Abstract: We present \textbf{Echo-Memory}, a controlled study of memory mechanisms in action-conditioned world models.
arXiv:2608.26671v1 Announce Type: new Abstract: Long autoregressive video generation faces a fundamental memory challenge: with a finite attention window, a model must decide which information from a...
arXiv:2607. 06481v1 Announce Type: cross Abstract: We present PACR-Video, a parameter-efficient framework for multi-shot long video extrapolation that preserves recurring entities, scene structure, visual style, and causal progression without full generator fine-tuning.
Short-drama generation has grown into a large, industrialized pipeline, and as it scales from isolated shots to the episode level, visual continuity has become a critical bottleneck. Current agent fra...
arXiv:2607. 15271v1 Announce Type: cross Abstract: Online novel view synthesis from multi-view streaming videos faces a fundamental trade-off: maintaining a persistent, long-horizon memory to reconstruct temporarily occluded regions while operating under strict real-time constraints.
arXiv:2606. 11792v1 Announce Type: cross Abstract: Video Large Multimodal Models have achieved remarkable progress in video understanding, yet they remain prone to hallucinations, where generated responses are not faithfully supported by the input video.
arXiv:2608.27280v1 Announce Type: new Abstract: Visual storytelling requires generating images that follow a narrative while preserving consistent character identities across frames. In free-form sto...
arXiv:2605. 21028v2 Announce Type: replace-cross Abstract: Autoregressive long video generation often adopts bounded-memory streaming for efficiency, typically combining local windows for short-term continuity with static early-frame sinks as long-range anchors.
arXiv:2606. 07577v1 Announce Type: new Abstract: Audio-visual large language models (LLMs) hold strong promise for long-form video understanding, yet their long-video inference is fundamentally limited by the linear growth of video tokens and key-value (KV) caches.
arXiv:2608.22869v1 Announce Type: cross Abstract: While Vision-Language-Action (VLA) models have leveraged internet-scale pretraining and task-focused finetuning to achieve strong performance on long...