The paper introduces a generalized multimodal foundation model that can handle arbitrary combinations of modalities and prediction tasks. It trains on large-scale synthetic multimodal datasets with diverse causal structures to learn transferable multimodal correlations. Experiments on 18 real-world datasets across 12 modalities and 11 tasks show competitive performance compared to specialized models without task-specific adaptation.
By Huizi Cui, Zongbo Han, Chenggong Ding, Naichuan Xiao, Jialong Yang, Jingdong Chen, Guangyu Wang, Qinghua Hu, Changqing Zhang
arXiv:2608.29335v1 Announce Type: new
Abstract: Latent generative models typically follow a two-stage pipeline, training a variational autoencoder for reconstruction and then a generative model on th...
By Guangting Zheng, Yiyuan Zhang, Tao Yang, Yunpeng Chen, Rui Zhu, Jiajun Deng, Yanyong Zhang
arXiv:2609.40362v1 Announce Type: new
Abstract: We present Multimodal Flow, a fully continuous generative model of language and vision. Most unified multimodal models either model both language and q...
By Hongyuan Tao, Xinggang Wang, Lianghui Zhu, Yongkang Li, Yunchao Wei, Bin Feng, Shaoyu Chen, Qian Zhang, Chang Huang, Kai Yu
AdaKerNet is a task‑adaptive neural kernel decoder that operates on frozen multimodal representations from large foundation models, without requiring access to the models’ parameters. It learns Lipschitz‑controlled multimodal features, a reference kernel providing a soft structural prior, and a lightweight nonlinear predictor that deforms this structure. Experiments on four multimodal large language models and diverse input modalities show consistent improvements over baseline decoders, achieving up to 41% error reduction in scarce‑label settings.
By Konstantinos D. Polyzos, Eleni Oikonomou, Tara Javidi
arXiv:2601. 06572v4 Announce Type: replace-cross Abstract: Multimodal variational autoencoders (VAEs) are widely used for weakly supervised generative learning with multiple modalities.
By Huyen Vo, Isabel Valera
arXiv:2608. 08135v1 Announce Type: cross Abstract: Cross-modality medical image translation can reduce the burden of multi-modal acquisitions, yet the field remains constrained by two coupled limitations: methods operate on 2D slices or 3D patches rather than whole volumes, and train a separate model for each translation task.
By Daniele Molino, Alessio Zoboli, Camillo Maria Caruso, Valerio Guarrasi, Paolo Soda