arXiv:2607. 24017v1 Announce Type: cross Abstract: The empirical success of attention mechanism in Multimodal Large Language Models (MLLMs) often obscures its inherent, subtle flaws.
By Pengkun Jiao, Bin Zhu, Jingjing Chen, Yu-gang Jiang
arXiv:2609.36776v1 Announce Type: new
Abstract: Emotional Video Captioning aims to generate factually accurate and emotionally empathetic descriptions. While recent methods have recognized the import...
By Cheng Ye, Weidong Chen, Peipei Song, Zhendong Mao
The study examines how Vision‑Language Models (VLMs) integrate visual evidence into language‑based decisions by applying layer‑wise causal interventions on video‑text attention pathways in a video‑based generative multiple‑choice setting. Findings reveal that visual information is primarily incorporated while processing candidate answer options, with nouns serving as key semantic anchors and verbs becoming important during temporal reasoning. The research also uncovers a distinct pattern in temporal reasoning, indicating that VLMs struggle to reconstruct sequential information across video frames, possibly due to linguistic biases in temporal expressions.
By Davide Testa, Hugh Mee Wong, Alessandro Lenci, Bernardo Magnini, Albert Gatt
arXiv:2606. 27596v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) exhibit sophisticated reasoning but remain susceptible to object hallucination.
By Liu Yu, Can Chen, Ping Kuang, Zhikun Feng, Fan Zhou, Gillian Dobbie
arXiv:2604. 15280v2 Announce Type: replace-cross Abstract: Understanding emotions is a fundamental ability for intelligent systems to be able to interact with humans.
By Madhav Agarwal, Sotirios A. Tsaftaris, Laura Sevilla-Lara, Steven McDonagh
The paper introduces DAN, a training‑free inference‑time framework that improves affective reasoning in multimodal large language models. It combines a Hierarchical Emotional Reasoning Chain (HERC) to better capture fine‑grained visual cues and a Contrastive Discriminative Visual Pruning (CDVP) module to isolate discriminative tokens for semantically similar emotions. Experiments show significant gains, notably a +10.47% improvement on the WebEmo25 benchmark with Qwen3‑VL‑8B‑Instruct.
By Cheng Ye, Weidong Chen, Zhaobo Qi, Beier Zhu, Zhendong Mao
arXiv:2608. 10448v1 Announce Type: new Abstract: Multimodal emotion recognition in conversation (MERC) requires understanding complex interactions between verbal and non-verbal cues.
By Sujung Oh, Jung Uk Kim, Sangmin Lee
arXiv:2607. 03358v1 Announce Type: cross Abstract: We study how visual information is routed in vision-language models (VLMs).
By Israfel Salazar, Stella Frank, Dan Oneata, Desmond Elliott, Constanza Fierro
arXiv:2609.36775v1 Announce Type: new
Abstract: Reinforcement Learning has significantly advanced the complex reasoning capabilities of MLLMs. However, prevailing RL algorithms suffer a severe failur...
By Cheng Ye, Weidong Chen, Bingyan Xu, Zhendong Mao
arXiv:2605. 16739v2 Announce Type: replace-cross Abstract: Decoding visual experience from brain activity has advanced substantially, but current brain-to-text systems largely recover semantic content while discarding affect.
By Bilal A. Mohammed, Lin Gu, Ruogu Fang
arXiv:2608. 03450v1 Announce Type: cross Abstract: Reasoning in Multimodal Large Language Models (MLLMs) requires both fine-grained visual perception and rigorous logical deduction.
By Haoqian Kang, Liupeng Li, Kuofeng Gao, Jinpeng Wang, Zhenyu Lu, Bin Chen, Ke Chen, Yaowei Wang
As generative multimedia evolves from static image synthesis to complex, interleaved visual narratives, a foundational bottleneck has emerged: the judgment crisis. While human perception naturally synthesizes the temporal and logical flow of a story, automated evaluation systems remain largely "blind" to sequential continuity, often failing to distinguish between a coherent narrative and a semantically shuffled or contradictory sequence.