VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2609.12818v1 Announce Type: new Abstract: Long video understanding often behaves like a visual needle-in-a-haystack problem: query-relevant evidence is sparsely distributed across long temporal...
arXiv:2609.09985v1 Announce Type: new Abstract: Real-world video understanding requires integrating visual, audio, textual, and temporal evidence distributed across a video. Yet many pipelines use a...
arXiv:2610.11171v1 Announce Type: cross Abstract: Long-form video understanding often involves multiple questions about different aspects of the same recording. Yet existing video agents typically pr...
arXiv:2607. 02588v1 Announce Type: cross Abstract: Multimodal large language models excel on short clips but struggle on hour-long videos in an online setting, where frames are processed incrementally under limited memory.
arXiv:2606. 07512v1 Announce Type: cross Abstract: Current Vision-Language Models struggle with hours-long videos because processing full-length visual sequences induces prohibitive token explosion and attention dilution.
arXiv:2608.31005v1 Announce Type: new Abstract: Existing long-video agents acquire evidence through one uniform behavior, ignoring whether the required evidence is concentrated, requires broad occurr...