arXiv Machine Learning

Efficient Multimodal Inference through Adaptive Acquisition and Sequential Fusion

arXiv AI
2d ago

Measure Less, Know More: Self-Supervised Test-Time Feature Acquisition

The paper introduces ECHO-$k$, a self-supervised, task-agnostic method for selecting which modalities to acquire at test time in multimodal, high-dimensional learning. By using a deep model’s pretrained representations as proxy targets, ECHO-$k$ learns a reinforcement‑learning policy that sequentially chooses informative modalities, providing theoretical guarantees in a linear setting. Experiments show that ECHO-$k$ consistently improves budgeted downstream performance across various foundation‑model backends, offering a principled approach to cost‑aware test‑time deployment when measurements are expensive or time‑constrained.

By Eeshaan Jain, Linus Bleistein, Bart Deplancke, Charlotte Bunne
arXiv Machine Learning
Jun 10

SynIB: Informational Bottleneck for Maximizing Synergy in Multimodal Learning

arXiv:2606. 09853v1 Announce Type: new Abstract: A central objective in multimodal learning is to capture synergy: task-relevant information that arises only from the joint use of multiple modalities, and is not available from any single modality alone.

By Konstantinos Kontras, Teodora Gagaleska, Thomas Strypsteen, Christos Chatzichristos, Matthew Blaschko, Maarten De Vos, Paul Pu Liang
arXiv AI
Jul 29

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities

arXiv:2607. 25948v1 Announce Type: cross Abstract: Any-to-any models predict any modality from any combination of others within a single network, a formulation used in multimodal vision and vision-language models, and increasingly in scientific domains such as ecology and astronomy.

By Mingqiao Ye, Zhaochong An, Zhitong Gao, Xian Liu, Fran\c{c}ois Fleuret, Chuan Li, Amir Zadeh, Serge Belongie, Afshin Dehghan, Jesse Allardice, David Mizrahi, O\u{g}uzhan Fatih Kar, Roman Bachmann, Amir Zamir
arXiv AI
Sep 30

AdaKerNet: Neural Kernel Decoding for Task-Adaptive Prediction with Multimodal Large Models

AdaKerNet is a task‑adaptive neural kernel decoder that operates on frozen multimodal representations from large foundation models, without requiring access to the models’ parameters. It learns Lipschitz‑controlled multimodal features, a reference kernel providing a soft structural prior, and a lightweight nonlinear predictor that deforms this structure. Experiments on four multimodal large language models and diverse input modalities show consistent improvements over baseline decoders, achieving up to 41% error reduction in scarce‑label settings.

By Konstantinos D. Polyzos, Eleni Oikonomou, Tara Javidi