arXiv AI By Yiheng Wang, Yueqian Lin, Lichen Zhu, Yudong Liu, Hai "Helen" Li, Yiran Chen

When No Answer Is Correct: Diagnosing Absent Answer Detection for MLLMs in Video Understanding

Read the original on arXiv AI →

arXiv:2606. 08239v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have made substantial advancements in video understanding, yet the reliability of their responses remains underexplored.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 2

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events

arXiv:2606. 02522v1 Announce Type: cross Abstract: Video multimodal large language models (MLLMs) have made rapid progress on general and long-form video understanding, yet their ability to preserve brief answer-critical visual evidence remains underexplored.

By Xiaolin Liu, Yilun Zhu, Xiangyu Zhao, Xuehui Wang, Yan Li, Xin Li, Haoyu Cao, Xing Sun, Shaofeng Zhang, Xu Yang, Zhihang Zhong, Xue Yang
arXiv Computation and Language
Sep 25

STRAND: Benchmarking and Improving Object-Centric Spatio-Temporal Monitoring in Video Large Language Models

STRAND is a new benchmark that tests multimodal large language models’ ability to track objects, their states, and relationships over time in videos. It evaluates intermediate reasoning by breaking queries into sub‑questions and uses Faithful Accuracy to ensure all parts of an answer are correct. The authors also propose an object‑centric framework that builds structured trajectories and shows reduced hallucinations and better temporal consistency compared to existing models.

By Thong Nguyen, Tri Cao, Khoi Le, Cong-Duy Nguyen, Quynh Vo, See-Kiong Ng, Bryan Hooi Kuen-Yew
arXiv Computation and Language
Sep 7

From Vision to Language: Investigating Causal Information Flow in Multimodal Decision-Making

The study examines how Vision‑Language Models (VLMs) integrate visual evidence into language‑based decisions by applying layer‑wise causal interventions on video‑text attention pathways in a video‑based generative multiple‑choice setting. Findings reveal that visual information is primarily incorporated while processing candidate answer options, with nouns serving as key semantic anchors and verbs becoming important during temporal reasoning. The research also uncovers a distinct pattern in temporal reasoning, indicating that VLMs struggle to reconstruct sequential information across video frames, possibly due to linguistic biases in temporal expressions.

By Davide Testa, Hugh Mee Wong, Alessandro Lenci, Bernardo Magnini, Albert Gatt
arXiv AI
Sep 2

TempCloze: Can Video-LLMs Identify the Missing Middle?

TempCloze is a video cloze benchmark designed to evaluate visual temporal reasoning in Video-LLMs. The task presents a video’s beginning and ending clips and asks models to select the correct missing middle from four candidates, focusing on semantic, alignment, and progression aspects while minimizing appearance cues. Evaluation of 31 models shows that temporal alignment is the main challenge, with models performing better on semantic content and event progression but struggling to place events correctly in time.

By Wenqi Pei, Henry Hengyuan Zhao, Yilai Liu, Jiahao Meng, Han Chen, Ziyu Wang, Hongyang Du
arXiv Computer Vision
Sep 16

Video-HolmesV2: Can MLLMs Reason with Spatio-Temporal Audio-Visual Evidence in Long Videos?

Video-HolmesV2 is a new benchmark that tests multimodal large language models on their ability to reason with spatio‑temporal audio‑visual evidence in long videos. It requires models to justify answers with precise evidence, uses a multi‑model cross‑verification pipeline and a spatio‑temporal evidence‑aware metric, and introduces an audio‑text guided token compression framework to reduce long‑context noise. In evaluations, even strong proprietary models score below 60% while the proposed approach outperforms comparable open‑source omni‑models.

By Zhaoyang Wei, Zipeng Wang, Yushe Cao, Chenhui Qiang, Shuaibing Cheng, Xuesong Yang, Sen Nie, Bowen Jiang, Wenchao Ding, Yanchao Hao, Zheng Wei, Xuehui Yu, Zhenjun Han