arXiv:2606. 00959v1 Announce Type: new Abstract: Understanding modality interaction in multimodal large language models (MLLMs) is central to reliable deployment.
By Wanlong Fang, Tianle Zhang, Wen Tao, Alvin Chan
arXiv:2608. 05381v1 Announce Type: new Abstract: Current Multimodal Large Language Models (MLLMs) can process diverse sensory inputs, yet their reasoning remains heavily biased toward a dominant modality, resulting in brittle cross-modal reasoning.
By Swapnanil Mukherjee, Agyeya Negi, Tanuja Ganu, Ponnurangam Kumaraguru
The paper introduces Inverted Asymmetric Fusion (IAF) to address strong-modality collapse in multimodal learning, where dominant modalities are degraded during fusion. IAF preserves the dominant modality by passing it unchanged and letting weaker modalities attend to it, while also strengthening weaker modalities via Modality-Aware Knowledge Distillation. Experiments on MultiHuSE, UR-FUNNY, and MUStARD show that IAF maintains unimodal performance and improves over the best unimodal baseline by up to 8.25%.
By Mary Ogbuka Kenneth, Foaad Khosmood, Abbas Edalat
arXiv:2603. 27958v2 Announce Type: replace Abstract: Analogical reasoning tests a fundamental aspect of human cognition: mapping the relation from one pair of objects to another.
By Yongkang Du, Xiaohan Zou, Minhao Cheng, Lu Lin
TwinICL is a procedurally generated benchmark that pairs matched text and image versions of tasks to enable controlled comparison of in‑context learning (ICL) across modalities. Experiments on six open‑weight models and 38 tasks show that multimodal ICL consistently underperforms text‑only ICL, with varying gaps by task family. Interventions targeting visual access, task framing, and reasoning can recover strong multimodal performance on a diagnostic subset, yet a modality gap remains even when explicit task instructions are provided, highlighting the dual role of demonstrations as context and evidence.
By Zihan Xue, Po-Yi Lu, Serhii Honcharenko, Zih-Ching Chen, Hsuan-Tien Lin, Nanyun Peng, I-Hung Hsu, Kuan-Hao Huang
arXiv:2606. 09853v1 Announce Type: new Abstract: A central objective in multimodal learning is to capture synergy: task-relevant information that arises only from the joint use of multiple modalities, and is not available from any single modality alone.
By Konstantinos Kontras, Teodora Gagaleska, Thomas Strypsteen, Christos Chatzichristos, Matthew Blaschko, Maarten De Vos, Paul Pu Liang