arXiv Machine Learning

Understanding Submodular Information Measure Based Objectives for Representation Learning: A Variance and Separation Perspective

arXiv:2607. 27660v1 Announce Type: new Abstract: Submodular Information Measures (SIMs) have recently emerged as a powerful framework for representation learning and multimodal learning.

arXiv Computer Vision
Sep 22

GeoBalance: Geometry-Aware Monitoring and Reconstruction with Asymmetric Optimization for Balanced Multimodal Learning

arXiv:2609.23533v1 Announce Type: new Abstract: Multimodal classifiers can converge to modality-dominant solutions in which one modality dominates the joint prediction, suppressing the learning of ot...

By Zechang Xiong, Da Li, Rong Yin, Kexin Tang, Biao Yang, Pengyuan Li, Wenkang Kong, Yulan Hu, Shengyu Zhu, Hao Peng
arXiv Computer Vision
Sep 16

Hub-Spectral Activation of Latent Multimodal Knowledge

arXiv:2609.17094v1 Announce Type: new Abstract: Multimodal representation learning seeks shared representations for cross-modal retrieval and knowledge transfer. Hub-based binding reduces pairwise su...

By Ying Guo, Haidong Chen, Linrui Xu, Xiaohao Liu, Chuancheng Shi, Canran Xiao, Dan Zhang, Fei Shen, Li Shen, Tat-Seng Chua
arXiv Machine Learning
Sep 10

Multi-Task Learning with Covariate-Overlap Regularization

The paper introduces COVER, a multi‑task learning framework that regularizes covariate overlap to mitigate the negative effects of sharing information across tasks with differing covariate distributions and response relationships. COVER blends a common component function, a shared neural representation, and low‑dimensional task‑specific coefficients, using taskwise second‑moment matrices to guide coefficient integration. The authors provide theoretical bias‑variance analysis, oracle inequalities, and neural‑network convergence rates, and demonstrate that COVER outperforms existing deep‑learning and statistical integration methods in simulations and a GTEx central‑nervous‑system study.

By Yang Sui, Qi Xu, Yang Bai, Annie Qu
arXiv Machine Learning
Jun 10

When to Align, When to Predict: A Phase Diagram for Multimodal Learning

arXiv:2606. 11190v1 Announce Type: new Abstract: Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds, when each fails, and when cross-modal training helps at all -- a gap that leaves practitioners, especially in scientific domains like biomedicine or astrophysics, with heterogeneous instruments and multiple levels of organization and measurement, unable to diagnose why standard methods underperform the best single modality.

By Ilay Kamai, Hugues Van Assel, Aviv Regev, Hagai B. Perets, Randall Balestriero
arXiv Computer Vision
3d ago

Mutual Equilibrium: Multimodal Representation Learning through Reciprocal Feedback

The paper introduces MEQ, a mutual feedback architecture that iteratively refines two multimodal inputs into coupled embeddings, each embedding incorporating information from the other. By continuously exchanging information between the modalities, the model converges to a fixed point that improves representation quality. Experiments on classification and visual grounding tasks show that MEQ achieves competitive or superior performance compared to concatenation-based baselines, and qualitatively enhances visual grounding when paired with complementary modalities.

By Ho-min Park, Byungkon Kang
Hugging Face Trending Papers
Jun 9

When to Align, When to Predict: A Phase Diagram for Multimodal Learning

Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds, when each fails, and when cross-modal training helps at all -- a gap that leaves practitioners, especially in scientific domains like biomedicine or astrophysics, with heterogeneous instruments and multiple levels of organization and measurement, unable to diagnose why standard methods underperform the best single modality. We develop a unified linear framework that addresses both questions.