arXiv Machine Learning By Mahir Shahriar Tamim, Sharjil Khan, Md. Samiul Alim, Tanvir Ahmed Khan, Shafin Rahman, Nabeel Mohammed

CAT-GS: Balanced Multimodal Learning via Calibrated Gating and Fusion Surgery

Read the original on arXiv Machine Learning →

The paper introduces CAT‑GS, a training controller that stabilizes multimodal neural networks by addressing three failure modes: modality imbalance, unstable gating, and fusion interference. CAT‑GS calibrates teacher-derived reliability, applies a margin‑thresholded gating policy, caps gradient budgets, and uses fusion‑only PCGrad, all without altering model architectures or losses. Experiments on audio‑visual, tri‑modal, synthetic, and cross‑domain benchmarks show that CAT‑GS matches or surpasses strong imbalance‑aware baselines while producing smoother gating and fewer conflicting fusion gradients.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 10

SynIB: Informational Bottleneck for Maximizing Synergy in Multimodal Learning

arXiv:2606. 09853v1 Announce Type: new Abstract: A central objective in multimodal learning is to capture synergy: task-relevant information that arises only from the joint use of multiple modalities, and is not available from any single modality alone.

By Konstantinos Kontras, Teodora Gagaleska, Thomas Strypsteen, Christos Chatzichristos, Matthew Blaschko, Maarten De Vos, Paul Pu Liang
arXiv Machine Learning
Jun 10

When to Align, When to Predict: A Phase Diagram for Multimodal Learning

arXiv:2606. 11190v1 Announce Type: new Abstract: Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds, when each fails, and when cross-modal training helps at all -- a gap that leaves practitioners, especially in scientific domains like biomedicine or astrophysics, with heterogeneous instruments and multiple levels of organization and measurement, unable to diagnose why standard methods underperform the best single modality.

By Ilay Kamai, Hugues Van Assel, Aviv Regev, Hagai B. Perets, Randall Balestriero