Multiplicity is an Inevitable and Inherent Challenge in Multimodal Learning
arXiv:2505. 19614v2 Announce Type: replace Abstract: Multimodal learning has seen remarkable progress, particularly with large-scale pre-training across various modalities.
The paper introduces a new multimodal scaling law that accounts for how the proportion of paired data influences loss in multimodal classification tasks. It shows that varying the pairing ratio under fixed total data budgets dramatically affects loss curves, with synergy only reducing when a critical threshold of paired data is reached. The proposed law decomposes total loss into four power-law components—redundancy, unique modality channels, and synergy—providing a more accurate fit (3.2% error) than existing models.
arXiv:2505. 19614v2 Announce Type: replace Abstract: Multimodal learning has seen remarkable progress, particularly with large-scale pre-training across various modalities.
The paper identifies that in multimodal learning, optimization often produces asymmetric certainty gains, with the stronger modality becoming more confident than the weaker one, which leads to imbalanced contributions and suboptimal performance. The authors attribute this issue to unimodal characteristics and propose a Max Confidence Regularization (MaxCR) method that tracks each modality’s semantic confidence via a nonlinear sparsity measure and applies max suppression and excitation to balance confidence levels. Experiments on standard datasets demonstrate that MaxCR improves overall performance compared to state‑of‑the‑art multimodal baselines.
arXiv:2606. 09853v1 Announce Type: new Abstract: A central objective in multimodal learning is to capture synergy: task-relevant information that arises only from the joint use of multiple modalities, and is not available from any single modality alone.
arXiv:2607. 05019v1 Announce Type: new Abstract: In multimodal classification, late-fusion approaches classify concatenated modality-specific features extracted by unimodal neural networks.
The paper investigates image tokenizers as the visual language of unified multimodal models by creating a controlled autoregressive testbed that tracks task‑specific validation losses during multimodal continual pretraining across text, image, text‑to‑image, and image‑to‑text predictions. It shows that losses must be analyzed by task, that the loss–performance relationship varies with the token space, and that better reconstruction does not always lead to stronger downstream performance. The study also demonstrates how tokenizer design choices—such as discriminator use, semantic supervision, and vocabulary size—affect joint modeling and downstream results.
arXiv:2606. 11190v1 Announce Type: new Abstract: Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds, when each fails, and when cross-modal training helps at all -- a gap that leaves practitioners, especially in scientific domains like biomedicine or astrophysics, with heterogeneous instruments and multiple levels of organization and measurement, unable to diagnose why standard methods underperform the best single modality.
arXiv:2606. 11614v1 Announce Type: cross Abstract: Multimodal learning hinges on capturing redundant, unique, and synergistic information across modalities, which collectively constitute multimodal interactions.
Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds, when each fails, and when cross-modal training helps at all -- a gap that leaves practitioners, especially in scientific domains like biomedicine or astrophysics, with heterogeneous instruments and multiple levels of organization and measurement, unable to diagnose why standard methods underperform the best single modality. We develop a unified linear framework that addresses both questions.
The paper introduces Inverted Asymmetric Fusion (IAF) to address strong-modality collapse in multimodal learning, where dominant modalities are degraded during fusion. IAF preserves the dominant modality by passing it unchanged and letting weaker modalities attend to it, while also strengthening weaker modalities via Modality-Aware Knowledge Distillation. Experiments on MultiHuSE, UR-FUNNY, and MUStARD show that IAF maintains unimodal performance and improves over the best unimodal baseline by up to 8.25%.
The paper introduces ModalShare, a bandwidth allocation method for multimodal split learning that assigns each modality a keep‑ratio based on its Shapley contribution score. Unlike existing compression schemes that split the uplink budget proportionally to activation size, ModalShare explicitly optimizes the split across modalities, requiring no extra uplink traffic or client computation. Experiments on CREMA‑D and MVSA datasets show that ModalShare improves accuracy by 12.4–15.4 percentage points over equal keep‑ratios under a 5× compression budget, outperforming three compressors across multiple datasets and budgets.
arXiv:2603. 13326v2 Announce Type: replace-cross Abstract: Multimodal Transformers often produce predictions without clarifying how different modalities jointly support a decision.
arXiv:2604. 18827v2 Announce Type: replace-cross Abstract: Scaling data and artificial neural networks has transformed AI, driving breakthroughs in language and vision.