arXiv Machine Learning By Zhou Du, Hamid Krim, Xiao Wu, Zhaoquan Yuan, Liangwei Li, Keisuke Fujii

Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA

Read the original on arXiv Machine Learning →

arXiv:2606. 09181v1 Announce Type: cross Abstract: Recent advances in video multimodal models have significantly improved VideoQA performance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computer Vision
Aug 21

Video Evidence to Reasoning Efficient Video Understanding via Explicit Evidence Grounding

arXiv:2601. 07761v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) face a fundamental dilemma in video reasoning: they are caught between the prohibitive computational costs of verbose reasoning and the hallucination risks of efficient, ungrounded approaches.

By Yanxiang Huang, Guohua Gao, Zhaoyang Wei
arXiv Computer Vision
6d ago

CCRV-Bench: Constraint-Based Evaluation of Causal Reasoning in Vision-Language Models

CCRV-Bench is a constraint‑driven benchmark designed to evaluate visual causal reasoning in vision‑language models on single‑image physical scenarios. It assesses four causal task dimensions—causal relation discovery, state prediction, causal diagnosis, and intervention—while applying constraints such as entity symbolization, spatial grounding, factual adversarial constraints, and minimalist output constraints to reduce shortcut learning. Experiments on 15 multimodal models reveal that constraint sensitivity varies by task and model, with intervention and spatial grounding having the largest impact and factual adversarial constraints improving causal diagnosis across models.

By Linyuan Gao, Yuan Wu, Yi Chang