arXiv AI

Evaluation of MLLM-Agnostic Plug-and-Play Keyframe Selection Methods for Long Video Understanding

The paper evaluates five training‑free, plug‑and‑play keyframe selection methods for multimodal large language models (MLLMs) on long‑video understanding tasks. It compares these methods across three different MLLMs and three video question‑answering benchmarks, finding that QAaF performs best in 13 of 15 settings while FOCUS ranks second. The study offers a unified benchmark for assessing MLLM‑agnostic keyframe selection techniques.

arXiv AI
Jul 29

Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding

arXiv:2607. 24794v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) demonstrate superior generalization in fundamental video tasks, restricted context windows limit their long video understanding.

By Linghao Meng, Qiankun Li, Junyuan Mao, Pujin Liao, Zhicheng He, Enbo Zhang, Kun Wang, Yang Liu, Huazhu Fu, Yueming Jin
arXiv AI
Jun 8

Watch, Remember, Reason: Human-View Video Understanding with MLLMs

arXiv:2606. 07433v1 Announce Type: cross Abstract: Video understanding is being rapidly transformed by multimodal large language models (MLLMs), as research moves from short clips to long, multimodal, and knowledge-intensive video scenarios.

By Jiahao Meng, Yue Tan, Qi Xu, Kuan Gao, Weisong Liu, Yanwei Li, Jason Li, Lingdong Kong, Haochen Wang, Qianyu Zhou, Jiangning Zhang, Guangliang Cheng, Yunhai Tong, Lu Qi, Minghsuan Yang
Hugging Face Trending Papers
Jul 2

ReQuest: Rethinking-based Question-Aware Frame Selection for Long-Form Video QA

Recent multimodal large language models (MLLMs) have substantially advanced video understanding, yet long-form video QA remains challenging under fixed input token budgets, where uniform sampling can be inefficient for evidence localization. We propose ReQuest , an uncertainty-driven, question-adaptive keyframe selection pipeline that aligns question intent with relevant video content through selective computation.

arXiv Computation and Language
Aug 31

Long Story Short: Story-level Video Understanding from 20K Short Films

The paper introduces Short‑Films 20K (SF20K), a large publicly available movie dataset comprising 20,143 amateur films totaling 3,582 hours, with an average length of 12 minutes per film. Accompanying the dataset is SF20K‑Test, a manual open‑ended question‑answering benchmark featuring 95 movies and 979 question‑answer pairs. Analysis of the benchmark shows limited data leakage, highlights the necessity of long‑term reasoning, and demonstrates that instruction tuning on the large‑scale dataset significantly boosts vision‑language model performance.

By Ridouane Ghermi, Xi Wang, Vicky Kalogeiton, Ivan Laptev
arXiv AI
Sep 10

Concord: A Video Relational Algebra for Cross-Modal Query Optimization

Concord introduces a Video Relational Algebra (VRA) that models videos, transcripts, frames, and object tracks, enabling semantic video queries. It applies approximate optimizations to rewrite VRA queries, reducing large language model (MLLM) usage by processing transcripts or using detection and tracking instead of full-video MLLM joins. Experiments on soccer broadcasts and lectures show that Concord sends only a small fraction of video to the MLLM, cutting costs by up to 87%, and improves cross‑camera query accuracy from an F1 of .364 to .813 without any MLLM calls.

By Sultan Muratbek, Charisse Ivana Yeung, Chanwut Kittivorawong, Alvin Cheung