arXiv AI By Jiyun Bae, Hyunjong Ok, Sangwoo Mo, Jaeho Lee

Understanding the Effects of Distractors on Reasoning Vision-Language Models

Read the original on arXiv AI →

arXiv:2511. 21397v2 Announce Type: replace-cross Abstract: How does irrelevant information (i.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 25

Mind What Matters for Reasoning: Aligning Cross-Modal Attention via Selective Probability Mass Concentration

The paper introduces Selective Probability Mass Concentration (sPMC), a training framework that strengthens implicit visual grounding in multimodal large language models by selectively regularizing attention heads most responsive to visual evidence. sPMC treats attention over visual tokens as a spatial probability distribution and encourages mass to concentrate on semantically relevant regions using segmentation-derived priors, while leaving other heads unconstrained. Across six multimodal benchmarks, sPMC yields an average zero‑shot improvement of 3% and gains up to 11.3% for various models by regularizing only 3%–15% of their attention heads.

By Jiaqi Deng, Zonghan Wu, Zhan Heng, Xiaoshui Huang, Huan Huo, Guandong Xu