arXiv AI

Disentangling Semantic Attention from Structural Bias in the Attention Manifold

arXiv:2607. 24017v1 Announce Type: cross Abstract: The empirical success of attention mechanism in Multimodal Large Language Models (MLLMs) often obscures its inherent, subtle flaws.

arXiv AI
Jun 24

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models

arXiv:2605. 08245v4 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) increasingly power high-stakes applications, from medical imaging to autonomous systems, yet they routinely hallucinate, confidently describing content not present in the input.

By Harshvardhan Saini, Samyak Jha, Yiming Tang, Dianbo Liu
arXiv AI
Sep 25

Mind What Matters for Reasoning: Aligning Cross-Modal Attention via Selective Probability Mass Concentration

The paper introduces Selective Probability Mass Concentration (sPMC), a training framework that strengthens implicit visual grounding in multimodal large language models by selectively regularizing attention heads most responsive to visual evidence. sPMC treats attention over visual tokens as a spatial probability distribution and encourages mass to concentrate on semantically relevant regions using segmentation-derived priors, while leaving other heads unconstrained. Across six multimodal benchmarks, sPMC yields an average zero‑shot improvement of 3% and gains up to 11.3% for various models by regularizing only 3%–15% of their attention heads.

By Jiaqi Deng, Zonghan Wu, Zhan Heng, Xiaoshui Huang, Huan Huo, Guandong Xu
arXiv AI
Sep 25

Interpreting and Enhancing Emotional Circuits in Large Vision-Language Models via Cross-Modal Information Flow

The paper introduces a steering‑vector‑based causal attribution framework to study how large vision‑language models (LVLMs) translate visual input into emotional narratives. By creating a specialized dataset, the authors uncover a functional decoupling in the LVLM’s three‑stage Adapt‑Aggregate‑Execute mechanism: visual emotional cues are first aggregated in middle layers via sentiment‑specific attention heads, then translated into narrative generation in deeper layers through emotion‑general pathways. Using these insights, they regulate emotional information routing to strengthen attention flow and amplify semantic activation, achieving significant performance gains on the MER‑UniBench and reducing emotional hallucinations through inference‑time intervention.

By Chengsheng Zhang, Chenghao Sun, Zhining Xie, Xinmei Tian
arXiv AI
Sep 2

Visual Attention Faithfulness in Vision-Language Models is Heterogeneous

The study investigates whether attention weights in Vision‑Language Models (VLMs) accurately reflect model reasoning for visual inputs. Using causal perturbation analysis, it identifies three distinct processing modes—Faithful‑Sufficient, Faithful‑Distributed, and Non‑Focal—indicating heterogeneous visual attention faithfulness. The research also shows that human‑annotated ground‑truth regions align with model attention in only about 60% of cases, highlighting a systematic divergence between model visual reliance and human intuition across VQA, document, and chart tasks.

By Xurui Song, Weishi Wang, Zhongqi Yue, Kuluhan Binici, Tao Bai, Hongxin Shao, Daniel Dahlmeier, Jun Luo
arXiv AI
Sep 10

STAR-Pro: Stage-Wise Token Adaptive Reduction with Progressive Refinement for Efficient Large Vision-Language Models

arXiv:2609.05916v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) achieve strong multimodal understanding, but the hundreds to thousands of visual tokens they process impose subs...

By Yichen Guo, Tinghao Wang, Qizhe Zhang, Lingbei Meng, Yuan Zhang, Jiajun Cao, Hao Jiang, Chenwei Wu, Jixian Wu, Sixiang Chen, Tao Luo, Hongyang Cheng, Kai Tang, Chenxi Li, Renyuan Li, Xiande Huang, Wenya Wang, Shanghang Zhang