arXiv:2606. 16408v1 Announce Type: new Abstract: We introduce MUNI, an end-to-end multimodal latent diffusion framework for any-to-any generation that unifies subset-conditioned cross-modal generation and unconditional joint sampling through a shared stochastic latent.
By Kyeongmin Yeo, Yunhong Min, Minhyuk Sung
arXiv:2608. 17203v1 Announce Type: cross Abstract: Contrastive learning has become a cornerstone of modern representation learning, powering CLIP-style models that underpin text-to-image generation, vision-language models, and retrieval across a rapidly growing range of modalities.
By Andrew Stuart, Florian Wolf
arXiv:2604. 07753v2 Announce Type: replace-cross Abstract: Empowering Large Multimodal Models (LMMs) with image generation often leads to catastrophic forgetting in understanding tasks due to severe gradient conflicts.
By Xiangyue Liu, Zijian Zhang, Miles Yang, Zhao Zhong, Liefeng Bo, Ping Tan
arXiv:2607. 05019v1 Announce Type: new Abstract: In multimodal classification, late-fusion approaches classify concatenated modality-specific features extracted by unimodal neural networks.
By Ilya Burenko, Dmitry Vetrov
arXiv:2606. 09853v1 Announce Type: new Abstract: A central objective in multimodal learning is to capture synergy: task-relevant information that arises only from the joint use of multiple modalities, and is not available from any single modality alone.
By Konstantinos Kontras, Teodora Gagaleska, Thomas Strypsteen, Christos Chatzichristos, Matthew Blaschko, Maarten De Vos, Paul Pu Liang
arXiv:2605. 05225v3 Announce Type: replace-cross Abstract: Mixture-of-Experts Multimodal Large Language Models (MoE MLLMs) suffer from a significant efficiency bottleneck during Expert Parallelism (EP) inference due to the straggler effect.
By Bo Li, Chuan Wu, Shaolin Zhu
arXiv:2608. 02769v1 Announce Type: cross Abstract: Multimodal supervised learning seeks to leverage multiple heterogeneous data sources to improve predictive performance.
By Sagnik Nandy, Samriddha Lahiry, Pragya Sur, Subhabrata Sen
arXiv:2606. 11614v1 Announce Type: cross Abstract: Multimodal learning hinges on capturing redundant, unique, and synergistic information across modalities, which collectively constitute multimodal interactions.
By Zequn Yang, Yake Wei, Haotian Ni, Zhihao Xu, Di Hu
arXiv:2505. 19614v2 Announce Type: replace Abstract: Multimodal learning has seen remarkable progress, particularly with large-scale pre-training across various modalities.
By Sanghyuk Chun, Olga Russakovsky
arXiv:2512. 04954v3 Announce Type: replace Abstract: We present a novel technique for amortized posterior estimation using Normalizing Flows trained with likelihood-weighted importance sampling.
By Rajneil Baruah
arXiv:2602. 17554v3 Announce Type: replace Abstract: Training large-scale generative models is resource-intensive and relies heavily on heuristic dataset weighting.
By Corinna Cortes, Mehryar Mohri, Yutao Zhong
arXiv:2607. 19332v1 Announce Type: new Abstract: Generative models have undergone many generations of evolution, from VAEs/GANs to diffusion/flow matching.
By Chirag Vashist, Ke Li