arXiv:2610.00573v1 Announce Type: new
Abstract: Query-aware keyframe selection enables multimodal large language models (MLLMs) to process long videos using only a small set of question-relevant fram...
By Haifeng Huang, Biyin Xu, Chunsheng Xin, Yang Li
arXiv:2604.17422v2 Announce Type: replace
Abstract: Long video understanding remains a formidable challenge for Multimodal Large Language Models (MLLMs) due to the prohibitive cost of processing dens...
By Shaoguang Wang, Weiyu Guo, Ziyang Chen, Xuming Hu, Hui Xiong
arXiv:2607. 24794v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) demonstrate superior generalization in fundamental video tasks, restricted context windows limit their long video understanding.
By Linghao Meng, Qiankun Li, Junyuan Mao, Pujin Liao, Zhicheng He, Enbo Zhang, Kun Wang, Yang Liu, Huazhu Fu, Yueming Jin
The paper evaluates five training‑free, plug‑and‑play keyframe selection methods for multimodal large language models (MLLMs) on long‑video understanding tasks. It compares these methods across three different MLLMs and three video question‑answering benchmarks, finding that QAaF performs best in 13 of 15 settings while FOCUS ranks second. The study offers a unified benchmark for assessing MLLM‑agnostic keyframe selection techniques.
By Dilip Sarkar, Md. Safayet Islam, Liang Liang
arXiv:2609.15408v1 Announce Type: cross
Abstract: Long-video understanding remains challenging for multimodal large language models (MLLMs) because densely encoding long frame sequences is computatio...
By Hongchang Shi, Jinpeng Hu, Ao Wang, Wenzheng Zhou, Hui Ma, Feng Li, Zenglin Shi
arXiv:2609.10008v1 Announce Type: new
Abstract: Composed video retrieval (CoVR) searches a gallery for the target video that realizes a natural-language modification of a source clip. However, at gal...
By Dmitry Demidov, Muhammad Zaigham Zaheer, Omkar Thawakar, Abdelrahman Mohamed Shaker, Rao Anwer