arXiv Computation and Language
Aug 31

Long Story Short: Story-level Video Understanding from 20K Short Films

The paper introduces Short‑Films 20K (SF20K), a large publicly available movie dataset comprising 20,143 amateur films totaling 3,582 hours, with an average length of 12 minutes per film. Accompanying the dataset is SF20K‑Test, a manual open‑ended question‑answering benchmark featuring 95 movies and 979 question‑answer pairs. Analysis of the benchmark shows limited data leakage, highlights the necessity of long‑term reasoning, and demonstrates that instruction tuning on the large‑scale dataset significantly boosts vision‑language model performance.

By Ridouane Ghermi, Xi Wang, Vicky Kalogeiton, Ivan Laptev
arXiv AI
Sep 10

TimeBlind: A Spatio-Temporal Compositionality Benchmark for Video LLMs

TimeBlind is a diagnostic benchmark designed to evaluate fine‑grained spatio‑temporal compositionality in video large language models (LLMs). It categorizes temporal understanding into three levels—atomic event recognition, event property characterization, and reasoning about event interdependencies—and uses a minimal‑pairs paradigm where video pairs share identical static content but differ only in temporal structure. Across 20 state‑of‑the‑art MLLMs tested on 600 curated instances, the best model achieved only 48.2% instance accuracy, far below human performance of 98.2%, highlighting a reliance on static visual shortcuts rather than true temporal reasoning.

By Baiqi Li, Kangyi Zhao, Ce Zhang, Chancharik Mitra, Jean de Dieu Nyandwi, Gedas Bertasius