arXiv:2607. 11078v1 Announce Type: cross Abstract: Can a Video Large Language Model (Video-LLM) follow one person through a long video, keeping track of who they are well enough to report, in order, how their outfit changes across a full TV episode?
By Mohammad Al-Ratrout, Shayla Sharmin, Aditya Raikwar, Roghayeh Leila Barmaki
arXiv:2607. 13305v1 Announce Type: cross Abstract: Benchmark accuracy in video large language models (LLMs) is often treated as evidence of visual understanding.
By Jae Joong Lee
arXiv:2606. 01485v1 Announce Type: cross Abstract: We describe our submission to the VRR Challenge @ CVPR 2026, built on the \emph{ImplicitQA} / \emph{VRR-QA} benchmark~\cite{implicitqa}: multiple-choice video question answering in which answers are deliberately \emph{not} observable in any single frame and must be inferred from spatial layout, motion, depth, viewpoint, causality, and social context across discontinuous frames of creative video.
By Ali Alavi
A score on a temporal video question answering benchmark is meant to measure that a model has temporal understanding, but it conflates two questions. 1.
arXiv:2607. 12304v1 Announce Type: cross Abstract: A score on a temporal video question answering benchmark is meant to measure that a model has temporal understanding, but it conflates two questions.
By Farrukh Rahman
arXiv:2608. 10385v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as relevance assessors in information retrieval (IR) evaluation, raising questions about how assessor framing affects judgment reliability and downstream system comparison.
By Samaneh Mohtadi, Pietro Bernardelle, Joel Mackenzie, Gianluca Demartini