arXiv:2605.26513v2 Announce Type: replace
Abstract: Even balanced multimodal learning methods do not consistently translate additional modalities into better regression performance. To understand thi...
By Haojie Yin, Chengcheng Feng, Tianyi Liu, Tianqi Zhang, Kaizhu Huang
The paper introduces CAT‑GS, a training controller that stabilizes multimodal neural networks by addressing three failure modes: modality imbalance, unstable gating, and fusion interference. CAT‑GS calibrates teacher-derived reliability, applies a margin‑thresholded gating policy, caps gradient budgets, and uses fusion‑only PCGrad, all without altering model architectures or losses. Experiments on audio‑visual, tri‑modal, synthetic, and cross‑domain benchmarks show that CAT‑GS matches or surpasses strong imbalance‑aware baselines while producing smoother gating and fewer conflicting fusion gradients.
By Mahir Shahriar Tamim, Sharjil Khan, Md. Samiul Alim, Tanvir Ahmed Khan, Shafin Rahman, Nabeel Mohammed
VCMM: Variance-Calibrated Momentum for Multimodal Learning proposes a new optimizer that adapts momentum based on modality-specific gradient dynamics. It estimates minibatch noise and temporal drift online, using a Kalman-inspired controller to set modality-specific momentum and applies bias correction for the first moment. Experiments on four multimodal benchmarks show consistent improvements with modest training overhead.
By Zhongjing Gu, Chenyang Huang, Yufa Feng, Chong He, Qinxu Ding, Yiming Cui
The paper identifies that in multimodal learning, optimization often produces asymmetric certainty gains, with the stronger modality becoming more confident than the weaker one, which leads to imbalanced contributions and suboptimal performance. The authors attribute this issue to unimodal characteristics and propose a Max Confidence Regularization (MaxCR) method that tracks each modality’s semantic confidence via a nonlinear sparsity measure and applies max suppression and excitation to balance confidence levels. Experiments on standard datasets demonstrate that MaxCR improves overall performance compared to state‑of‑the‑art multimodal baselines.
By Longfei Huang, Xiangyu Wu, Yang Yang
arXiv:2609.16059v1 Announce Type: cross
Abstract: Multimodal instruction following (MMIF) is crucial for building generalist agents. However, current training paradigms rely heavily on Supervised Fin...
By Yirong Zeng, Zhang Sai, Yuxian Wang, Yutai Hou, Yufei Liu, Xiao Ding, Bibo Cai
Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds, when each fails, and when cross-modal training helps at all -- a gap that leaves practitioners, especially in scientific domains like biomedicine or astrophysics, with heterogeneous instruments and multiple levels of organization and measurement, unable to diagnose why standard methods underperform the best single modality. We develop a unified linear framework that addresses both questions.
arXiv:2606. 09853v1 Announce Type: new Abstract: A central objective in multimodal learning is to capture synergy: task-relevant information that arises only from the joint use of multiple modalities, and is not available from any single modality alone.
By Konstantinos Kontras, Teodora Gagaleska, Thomas Strypsteen, Christos Chatzichristos, Matthew Blaschko, Maarten De Vos, Paul Pu Liang
arXiv:2609.23533v1 Announce Type: new
Abstract: Multimodal classifiers can converge to modality-dominant solutions in which one modality dominates the joint prediction, suppressing the learning of ot...
By Zechang Xiong, Da Li, Rong Yin, Kexin Tang, Biao Yang, Pengyuan Li, Wenkang Kong, Yulan Hu, Shengyu Zhu, Hao Peng
arXiv:2606. 11190v1 Announce Type: new Abstract: Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds, when each fails, and when cross-modal training helps at all -- a gap that leaves practitioners, especially in scientific domains like biomedicine or astrophysics, with heterogeneous instruments and multiple levels of organization and measurement, unable to diagnose why standard methods underperform the best single modality.
By Ilay Kamai, Hugues Van Assel, Aviv Regev, Hagai B. Perets, Randall Balestriero
arXiv:2607. 05019v1 Announce Type: new Abstract: In multimodal classification, late-fusion approaches classify concatenated modality-specific features extracted by unimodal neural networks.
By Ilya Burenko, Dmitry Vetrov
arXiv:2607. 16789v1 Announce Type: new Abstract: Real-world perception and decision making are inherently multimodal, integrating complementary signals across modalities.
By Sana Tonekaboni, Viktoria Schuster, Caroline Uhler
arXiv:2609.25836v1 Announce Type: new
Abstract: Multi-task optimization (MTO) addresses a set of optimization tasks simultaneously, often suffering from inaccurate inter-task relationship estimation...
By Tingyang Wei, Haofeng Wu, Jiao Liu, Zhao Wei, Puay Siew Tan, Yew-Soon Ong