The paper introduces TRACE, a training‑free agent for long‑video understanding that grounds answers in raw visual clips and builds an evidence bundle until the answer stabilises. It also presents VES‑Bench, a 600‑question benchmark over 348 public long videos that audits whether decoded frames cover all necessary evidence intervals at three strictness levels. TRACE achieves high accuracy on VES‑Bench (63.5% audit accuracy) and remains competitive on other video‑understanding benchmarks while using far fewer frames than uniform decoding.
By Pengyiang Liu, Junbo Niu, Xiaoyang Hu, Zhongyue Shi, Zitian Wang, Linjiang Huang, Si Liu
arXiv:2605.17610v2 Announce Type: replace-cross
Abstract: The rapid growth of online video platforms and AI-generated content has made reliable video guardrails a key challenge for safety and real-wo...
By Shahriar Kabir Nahin, Hadi Askari, Muhao Chen, Anshuman Chhabra
arXiv:2609.12818v1 Announce Type: new
Abstract: Long video understanding often behaves like a visual needle-in-a-haystack problem: query-relevant evidence is sparsely distributed across long temporal...
By Sen Yang, Boqiang Duan, Jing Yang, Weihao Bo, Jie Liu, Boyuan Tong, Ze Feng, Wenkang Zhang, Jingdong Wang, Hua Wu
arXiv:2608.23329v1 Announce Type: cross
Abstract: Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from the video a...
By Wenqi Liu, Shijie Ma, Yunxiao Wang, Meng Liu, Qile Su, Han Liu, Bohan Hou, Xuanyu Zheng, Changyi Liu, Tianke Zhang, Haonan Fan, Kaiyu Jiang, Yingxin Li, Jiankang Chen, Xu Wang, Bin Wen, Tingting Gao, Han Li, Jianhua Yin, Yinwei Wei, Xuemeng Song
arXiv:2509.01167v3 Announce Type: replace-cross
Abstract: Vision-language models (VLMs) can ingest only a limited number of video frames, making frame selection a practical necessity. But do current...
By Hyunjong Ok, Jaeho Lee
arXiv:2609.38413v1 Announce Type: new
Abstract: Vision-language models (VLMs) can answer questions about hour-long videos, but processing every frame is prohibitively expensive, even though the evide...
By Susan Liang, Jianmin Wu, Daxiang Dong
The paper introduces CoVR‑R, a reason‑aware composed video retrieval system that, given a reference video and an edit instruction, retrieves a target video that satisfies the edit. It employs a zero‑shot reason‑then‑retrieve pipeline using Qwen3.5‑27B to generate structured descriptions and dense embeddings for gallery videos, and performs edit reasoning on the query to produce a target‑video description used as the query embedding. The method combines dense retrieval with a TF‑IDF branch over generated texts, fusing the rankings with split‑specific weights, achieving state‑of‑the‑art retrieval metrics on both validation and blind test splits.
By Dongqing Liu, Mengshi Qi, Hongwei Ji
arXiv:2606. 13141v1 Announce Type: new Abstract: Retrieval-augmented generation is moving beyond text into long, egocentric video, where systems must select query-relevant chunks across multiple modalities and temporal granularities.
By Yuho Lee, Jisu Shin, Nicole Hee-Yeon Kim, Jihwan Bang, Juntae Lee, Kyuwoong Hwang, Fatih Porikli, Hwanjun Song
arXiv:2607. 17279v1 Announce Type: cross Abstract: Recently, text-to-video (T2V) models have been widely deployed, sparking growing concerns over their robustness against jailbreak attacks.
By Xingkai Peng, Jun Jiang, Jiayang Liu, Kejiang Chen, Weiming Zhang
The study investigates how long‑video language models decide which frames to keep, compress, and reuse, testing each decision in isolation across six selection rules, three benchmarks, and two answering models. It finds that selecting frames based on queries yields the biggest performance boost, that halving spatial resolution costs little, and that reallocating saved tokens to more compressed frames can further improve accuracy. The work also highlights the importance of a unified evaluation harness to avoid misleading comparisons.
By Prakhar Khatri
arXiv:2606. 26904v1 Announce Type: cross Abstract: Video reasoning language models implicitly assume that every input frame is equally reliable.
By Yangfan He, Yujin Choi, Jaehong Yoon
Recent advancements in Image-to-Video (I2V) generation have transformed input images from simple appearance references into interactive control interfaces where visual cues such as arrows, sketches, and emojis orchestrate complex video dynamics with unprecedented controllability. However, these seemingly innocuous static cues can be interpreted by models as executable temporal instructions, unfolding into harmful actions in the generated videos.