arXiv Machine Learning

CAT-GS: Balanced Multimodal Learning via Calibrated Gating and Fusion Surgery

The paper introduces CAT‑GS, a training controller that stabilizes multimodal neural networks by addressing three failure modes: modality imbalance, unstable gating, and fusion interference. CAT‑GS calibrates teacher-derived reliability, applies a margin‑thresholded gating policy, caps gradient budgets, and uses fusion‑only PCGrad, all without altering model architectures or losses. Experiments on audio‑visual, tri‑modal, synthetic, and cross‑domain benchmarks show that CAT‑GS matches or surpasses strong imbalance‑aware baselines while producing smoother gating and fewer conflicting fusion gradients.

arXiv Machine Learning
Jun 10

SynIB: Informational Bottleneck for Maximizing Synergy in Multimodal Learning

arXiv:2606. 09853v1 Announce Type: new Abstract: A central objective in multimodal learning is to capture synergy: task-relevant information that arises only from the joint use of multiple modalities, and is not available from any single modality alone.

By Konstantinos Kontras, Teodora Gagaleska, Thomas Strypsteen, Christos Chatzichristos, Matthew Blaschko, Maarten De Vos, Paul Pu Liang
arXiv Machine Learning
Jun 10

When to Align, When to Predict: A Phase Diagram for Multimodal Learning

arXiv:2606. 11190v1 Announce Type: new Abstract: Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds, when each fails, and when cross-modal training helps at all -- a gap that leaves practitioners, especially in scientific domains like biomedicine or astrophysics, with heterogeneous instruments and multiple levels of organization and measurement, unable to diagnose why standard methods underperform the best single modality.

By Ilay Kamai, Hugues Van Assel, Aviv Regev, Hagai B. Perets, Randall Balestriero
Hugging Face Trending Papers
Jun 9

When to Align, When to Predict: A Phase Diagram for Multimodal Learning

Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds, when each fails, and when cross-modal training helps at all -- a gap that leaves practitioners, especially in scientific domains like biomedicine or astrophysics, with heterogeneous instruments and multiple levels of organization and measurement, unable to diagnose why standard methods underperform the best single modality. We develop a unified linear framework that addresses both questions.

arXiv AI
Aug 25

Mitigating Sample-Level Imbalance via Probabilistic Separation for Adaptive Multimodal Fusion

The paper introduces a framework to tackle modality imbalance in multimodal learning by focusing on sample-level variations. It defines a Modality Gap metric to measure prediction discrepancies, models the resulting bimodal distribution with a Gaussian Mixture Model, and uses Bayesian probabilities for soft separation of balanced and imbalanced samples. A two‑stage training process—Warm‑up and Adaptive Training—reallocates loss weights based on the GMM, strengthening alignment for imbalanced samples while favoring fusion for balanced ones, and shows superior performance over existing baselines.

By Zhiwen Yu, Zhaocheng Liu, Xiaoqing Liu, Huanqiang Zeng, C. L. Philip Chen
arXiv AI
Aug 18

The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning

The Unwritten Benchmark introduces a novel challenge for multimodal machine learning, focusing on abstract perceptual reasoning through acousto‑kinematic word inference. Models must decode words written only by the audio of pen scratches and the video of hand movements, across three writing styles, without any visible ink trace. Evaluation shows a stark performance gap: humans achieve over 80% ordered letter accuracy, while leading models like GPT‑4o and Gemini 2.5‑Pro fail to exceed 10%, and combining modalities often degrades performance.

By Garima Arya Yadav, Nilay Yilmaz, Yezhou Yang