The paper introduces DAN, a training‑free inference‑time framework that improves affective reasoning in multimodal large language models. It combines a Hierarchical Emotional Reasoning Chain (HERC) to better capture fine‑grained visual cues and a Contrastive Discriminative Visual Pruning (CDVP) module to isolate discriminative tokens for semantically similar emotions. Experiments show significant gains, notably a +10.47% improvement on the WebEmo25 benchmark with Qwen3‑VL‑8B‑Instruct.
By Cheng Ye, Weidong Chen, Zhaobo Qi, Beier Zhu, Zhendong Mao
arXiv:2602. 23802v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual reasoning and understanding tasks but still struggle to capture the complexity and subjectivity of human emotions.
By Yiyang Fang, Wenke Huang, Pei Fu, Yihao Yang, Kehua Su, Zhenbo Luo, Jian Luan, Mang Ye
The paper introduces a steering‑vector‑based causal attribution framework to study how large vision‑language models (LVLMs) translate visual input into emotional narratives. By creating a specialized dataset, the authors uncover a functional decoupling in the LVLM’s three‑stage Adapt‑Aggregate‑Execute mechanism: visual emotional cues are first aggregated in middle layers via sentiment‑specific attention heads, then translated into narrative generation in deeper layers through emotion‑general pathways. Using these insights, they regulate emotional information routing to strengthen attention flow and amplify semantic activation, achieving significant performance gains on the MER‑UniBench and reducing emotional hallucinations through inference‑time intervention.
By Chengsheng Zhang, Chenghao Sun, Zhining Xie, Xinmei Tian
arXiv:2609.40079v1 Announce Type: cross
Abstract: While recent Multimodal Large Language Models (MLLMs) have shown promise in affective computing, their reasoning capabilities are largely confined to...
By Shuo Zhang, Yifan Zhou, Han Wang, Jinsong Zhang, Jingyu Li, Hongbing Li, Zhejun Zhang, Chengyi Zhao, Yuquan Hao, Yitong Liu, Jiyin Li, Ruiqi Tang, Zixuan Lin, Yi Luo, Xurui Zhang, Ronghao Chen, Huacan Wang, Lei Li
arXiv:2604. 15280v2 Announce Type: replace-cross Abstract: Understanding emotions is a fundamental ability for intelligent systems to be able to interact with humans.
By Madhav Agarwal, Sotirios A. Tsaftaris, Laura Sevilla-Lara, Steven McDonagh
arXiv:2607. 10678v1 Announce Type: new Abstract: Emotional intelligence enables humans to recognize emotions, infer their causes, reason about interventions, and modify their environment to achieve desired affective states.
By Qing Lin, Mengmi Zhang
arXiv:2608.21022v1 Announce Type: new
Abstract: Micro-actions are subtle, short, low-amplitude body movements, such as a fidgeting hand or a slight head tilt, that humans perform with little consciou...
By Fengshun Wang, Jin'ang Han, Zhigang Tu
The paper introduces ProAction, a multimodal dataset of 10,000 samples comprising visual, audio, and text inputs across 12 daily-life scenarios, designed to support the Proactive Robot Action Reasoning (ProRobo) problem. It presents a two-stage human-in-the-loop annotation pipeline that incorporates appraisal and Theory-of-Mind considerations to generate cognitively grounded high-level action labels. The authors benchmark multimodal large language models and propose MMC2Act, showing that training on ProAction significantly improves proactive action reasoning compared to general-purpose models.
By Zhihao Gu, Kechao Zhu, Yuanfeng Wu, Mohan Liu, Ankit Kumar Shaw, ChenDong Hong, Xuanyu Chen, Dengchen Mei, Xu Tianyi, Lin Wang
arXiv:2510. 02528v2 Announce Type: replace Abstract: Large Multimodal Models (LMMs) demonstrate impressive in-context learning abilities from few multimodal demonstrations, yet the internal mechanisms supporting such task learning remain opaque.
By Shuhao Fu, Esther Goldberg, Ying Nian Wu, Hongjing Lu
arXiv:2606. 27652v1 Announce Type: new Abstract: We find that explicit reasoning does not necessarily translate into better multimodal emotion recognition (MER) accuracy, even though it makes predictions more interpretable.
By Zhiyuan Han, Beier Zhu, Wenwen Tong, Chengwei Qin, Xinyi Wang, Jiayu Zhang, Jiangnan Chen, Hewei Guo, Dongchuan Ran, Lewei Lu, Xun Yang
The paper introduces LIFT, a lightweight vector‑intervention technique that transfers reasoning capability from a base large language model (LLM) to a vision‑language model (VLM) without retraining the VLM backbone. LIFT defines Reasoning Vectors as differences in hidden states between a reasoning path with an explicit trace and a solver path without it, and injects these vectors into the VLM’s language‑side activations. Experiments on two VLMs across six reasoning benchmarks show that vectors derived from the base LLM consistently outperform those derived from the aligned VLM, indicating that the base LLM is a more effective source for recovering degraded reasoning.
"whyItMatters":"The study demonstrates that a simple, frozen‑backbone intervention can partially restore reasoning abilities in multimodal models, highlighting the value of leveraging the original language model’s reasoning power."
By Ziyi Wang, Li Li, Aolin Zhou, Yankun Shen, Chonghan Liu, Shuxia Lin, Xu Yang
arXiv:2511. 17731v2 Announce Type: replace-cross Abstract: Chain-of-Thought (CoT) prompting has proven remarkably effective for eliciting complex reasoning in large language models (LLMs).
By Lingxiao Li, Yifan Wang, Xinyan Gao, Chen Tang, Xiangyu Yue, Chenyu You