arXiv AI

Mechanistic Insights into Functional Sparsity in Multimodal LLMs via CoRe Heads

arXiv:2606. 05843v1 Announce Type: cross Abstract: While Multimodal Large Language Models (MLLMs) demonstrate remarkable proficiency on complex vision-language tasks, the mechanisms by which they extract query-relevant visual features from complex, noisy contexts remain opaque.