SportsGrounder: Proposal-Aided Interleaved Grounding for Dense Sports Video Reasoning
arXiv:2608. 07932v2 Announce Type: replace Abstract: Sports video analysis is crucial for athletic analytics and broadcasting enhancement.
arXiv:2606. 09181v1 Announce Type: cross Abstract: Recent advances in video multimodal models have significantly improved VideoQA performance.
arXiv:2608. 07932v2 Announce Type: replace Abstract: Sports video analysis is crucial for athletic analytics and broadcasting enhancement.
arXiv:2601. 07761v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) face a fundamental dilemma in video reasoning: they are caught between the prohibitive computational costs of verbose reasoning and the hallucination risks of efficient, ungrounded approaches.
arXiv:2609.36776v1 Announce Type: new Abstract: Emotional Video Captioning aims to generate factually accurate and emotionally empathetic descriptions. While recent methods have recognized the import...
CCRV-Bench is a constraint‑driven benchmark designed to evaluate visual causal reasoning in vision‑language models on single‑image physical scenarios. It assesses four causal task dimensions—causal relation discovery, state prediction, causal diagnosis, and intervention—while applying constraints such as entity symbolization, spatial grounding, factual adversarial constraints, and minimalist output constraints to reduce shortcut learning. Experiments on 15 multimodal models reveal that constraint sensitivity varies by task and model, with intervention and spatial grounding having the largest impact and factual adversarial constraints improving causal diagnosis across models.
arXiv:2508. 21010v3 Announce Type: replace-cross Abstract: Existing Causal-Why Video Question Answering (VideoQA) models often struggle with higher-order reasoning, relying on opaque, monolithic pipelines that entangle video understanding, causal inference, and answer generation.
arXiv:2607. 08763v1 Announce Type: cross Abstract: Reasoning has become a core capability for large models, especially when reliable decisions require understanding logical consequences.
The paper introduces the Very Big Video Reasoning (VBVR) Dataset, a large-scale collection of over one million video clips organized into 200 curated reasoning tasks. It also presents VBVR-Bench, a benchmark framework that uses rule-based, human-aligned scorers for reproducible evaluation of video reasoning models. The authors conduct a large-scale scaling study, noting early signs of emergent generalization to unseen reasoning tasks, and make all resources publicly available.
arXiv:2608.23330v1 Announce Type: new Abstract: Video understanding requires intelligent agents to transcend mere recognition of visual facts and comprehend the underlying intents behind human action...
arXiv:2607. 21267v1 Announce Type: new Abstract: Comprehensive basketball video understanding requires resolving not only what event occurs, but also who is responsible and when the key evidence appears.
arXiv:2608. 06938v1 Announce Type: cross Abstract: The visual reasoning ability of multimodal large language models (MLLMs) is crucial for downstream applications, particularly counter-commonsense reasoning, which requires models to reason beyond common assumptions.
arXiv:2505.01583v2 Announce Type: replace Abstract: Understanding causal event relationships and achieving fine-grained temporal grounding in videos remain challenging for vision-language models (VLM...
arXiv:2606. 07433v1 Announce Type: cross Abstract: Video understanding is being rapidly transformed by multimodal large language models (MLLMs), as research moves from short clips to long, multimodal, and knowledge-intensive video scenarios.