arXiv:2607. 08839v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are typically designed under the assumption that all modalities available during training will also be accessible at inference.
By Dominick Reilly, Qiyu Wu, Hiromi Wakaki, Srijan Das, Yuki Mistufuji
arXiv:2610.07499v1 Announce Type: new
Abstract: Multimodal wearable systems must remain reliable when sensor streams become noisy or unavailable. Existing multimodal test-time adaptation (TTA) method...
By Payal Mohapatra, Yueyuan Sui, Haodong Yang, Benjamin Lundell, Stephen Xia, Qi Zhu
The paper introduces ECHO-$k$, a self-supervised, task-agnostic method for selecting which modalities to acquire at test time in multimodal, high-dimensional learning. By using a deep model’s pretrained representations as proxy targets, ECHO-$k$ learns a reinforcement‑learning policy that sequentially chooses informative modalities, providing theoretical guarantees in a linear setting. Experiments show that ECHO-$k$ consistently improves budgeted downstream performance across various foundation‑model backends, offering a principled approach to cost‑aware test‑time deployment when measurements are expensive or time‑constrained.
By Eeshaan Jain, Linus Bleistein, Bart Deplancke, Charlotte Bunne
arXiv:2606. 09853v1 Announce Type: new Abstract: A central objective in multimodal learning is to capture synergy: task-relevant information that arises only from the joint use of multiple modalities, and is not available from any single modality alone.
By Konstantinos Kontras, Teodora Gagaleska, Thomas Strypsteen, Christos Chatzichristos, Matthew Blaschko, Maarten De Vos, Paul Pu Liang
arXiv:2608. 02769v1 Announce Type: cross Abstract: Multimodal supervised learning seeks to leverage multiple heterogeneous data sources to improve predictive performance.
By Sagnik Nandy, Samriddha Lahiry, Pragya Sur, Subhabrata Sen
arXiv:2608. 05608v1 Announce Type: cross Abstract: Multimodal classification typically assumes all modalities are available, yet real-world inputs are often incomplete.
By Yunping Shi, En Yu, Kairui Guo, Jie Lu
arXiv:2607. 25948v1 Announce Type: cross Abstract: Any-to-any models predict any modality from any combination of others within a single network, a formulation used in multimodal vision and vision-language models, and increasingly in scientific domains such as ecology and astronomy.
By Mingqiao Ye, Zhaochong An, Zhitong Gao, Xian Liu, Fran\c{c}ois Fleuret, Chuan Li, Amir Zadeh, Serge Belongie, Afshin Dehghan, Jesse Allardice, David Mizrahi, O\u{g}uzhan Fatih Kar, Roman Bachmann, Amir Zamir
arXiv:2609.16059v1 Announce Type: cross
Abstract: Multimodal instruction following (MMIF) is crucial for building generalist agents. However, current training paradigms rely heavily on Supervised Fin...
By Yirong Zeng, Zhang Sai, Yuxian Wang, Yutai Hou, Yufei Liu, Xiao Ding, Bibo Cai
AdaKerNet is a task‑adaptive neural kernel decoder that operates on frozen multimodal representations from large foundation models, without requiring access to the models’ parameters. It learns Lipschitz‑controlled multimodal features, a reference kernel providing a soft structural prior, and a lightweight nonlinear predictor that deforms this structure. Experiments on four multimodal large language models and diverse input modalities show consistent improvements over baseline decoders, achieving up to 41% error reduction in scarce‑label settings.
By Konstantinos D. Polyzos, Eleni Oikonomou, Tara Javidi
arXiv:2606. 02679v1 Announce Type: new Abstract: Multimodal systems often benefit from combining information across language, sound, and visual streams, but this benefit is not guaranteed.
By Jiyuan Liu, Liangwei Nathan Zheng, Wei Emma Zhang, Xinpei Wang, Weitong Chen
arXiv:2607. 27766v1 Announce Type: cross Abstract: On-device in-context learning (ICL) relies on pre-inference retrieval to select demonstrations for useful context before downstream model inference.
By Xinyu Luo, Hui Liu, Yihua Shao, Junyi Yang, Arindam Basu, Haoliang Li
arXiv:2506. 01850v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable success in instruction-following tasks by integrating pretrained visual encoders with large language models (LLMs).
By Wayner Barrios, Andr\'es Villa, Juan Le\'on Alc\'azar, SouYoung Jin, Bernard Ghanem