arXiv:2601. 22574v2 Announce Type: replace-cross Abstract: Although Video Large Multimodal Models have achieved strong performance in video understanding, they still suffer from hallucination.
By Yuansheng Gao, Jinman Zhao, Tong Zhang, Xingguo Xu, Wenbin Xing, Han Bao, Zonghui Wang, Wenzhi Chen
CounterVid introduces a scalable counterfactual video generation framework that creates videos differing only in actions or temporal structure while keeping scene context intact. The approach uses multimodal LLMs for action proposals and diffusion models for editing, producing a synthetic dataset of ~26k preference pairs for action recognition and sequence ordering. With the MixDPO optimization method, the authors demonstrate significant improvements in action recognition and temporal ordering on Qwen2.5‑VL and InternVL3 backbones, while maintaining overall video understanding.
By Tobia Poppi, Burak Uzkent, Amanmeet Garg, Lucas Porto, Garin Kessler, Yezhou Yang, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara, Florian Schiffers
arXiv:2609.16646v1 Announce Type: new
Abstract: When strong multimodal models are widely available, progress requires new scientific methodologies beyond benchmark scores---using models as instrument...
By Zhipeng Zhao, Wenxu Wang, Peishun Liu, Ruichun Tang
arXiv:2609.36440v1 Announce Type: new
Abstract: Large vision-language models (LVLMs) have recently achieved remarkable progress across multimodal tasks, yet object hallucination remains a persistent...
By Jae-Ho Lee, Jeong-Eun Lee, Gyeong-Moon Park
The paper examines how hallucinations arise in multi-stage video‑understanding agents by aligning existing benchmarks with the stages of temporal grounding, visual observation, and reasoning. It introduces a causal stage‑intervention protocol that isolates each stage while keeping the downstream task constant, revealing that grounding errors dominate downstream hallucinations and that correct region location matters more than precise temporal overlap. The study also shows that current benchmark scores poorly predict causal sensitivity and can fail under distribution shift, advocating for stage‑aware evaluation methods.
By Shuzhi Gong, Fengze Sun, Yuansan Liu
arXiv:2607. 07507v1 Announce Type: cross Abstract: Hallucinations in vision language models (VLMs) are commonly treated as semantic errors, yet they often arise from partial or ambiguous visual evidence.
By Feng He, Zhenting Wang, Qifan Wang, Qiang Guan, Dongfang Liu, Ruixiang Tang, Qiankun Li