Multimodality as Supervision: Self-Supervised Specialization to the Test Environment via Multimodality
Cross-modal learning, i. e.
The paper introduces ECHO-$k$, a self-supervised, task-agnostic method for selecting which modalities to acquire at test time in multimodal, high-dimensional learning. By using a deep model’s pretrained representations as proxy targets, ECHO-$k$ learns a reinforcement‑learning policy that sequentially chooses informative modalities, providing theoretical guarantees in a linear setting. Experiments show that ECHO-$k$ consistently improves budgeted downstream performance across various foundation‑model backends, offering a principled approach to cost‑aware test‑time deployment when measurements are expensive or time‑constrained.
Cross-modal learning, i. e.
arXiv:2607. 14721v1 Announce Type: cross Abstract: Cross-modal learning, i.
arXiv:2607. 08839v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are typically designed under the assumption that all modalities available during training will also be accessible at inference.
arXiv:2510. 12624v2 Announce Type: replace-cross Abstract: Active feature acquisition (AFA) is a sequential decision-making problem where the goal is to improve model performance for test instances by adaptively selecting which features to acquire.
The paper introduces a generalized multimodal foundation model that can handle arbitrary combinations of modalities and prediction tasks. It trains on large-scale synthetic multimodal datasets with diverse causal structures to learn transferable multimodal correlations. Experiments on 18 real-world datasets across 12 modalities and 11 tasks show competitive performance compared to specialized models without task-specific adaptation.
arXiv:2608. 12724v1 Announce Type: new Abstract: Few-shot in-context learning (ICL) with multi-modal large language models (MLLMs) enables task adaptation without parameter updates, but its performance is highly sensitive to the quality and coverage of the selected demonstrations.
arXiv:2608. 01672v1 Announce Type: cross Abstract: Effective long-context modeling is not merely about retaining more of the past, but about preserving the information that may prove relevant later.
arXiv:2604. 01577v3 Announce Type: replace-cross Abstract: We study out of distribution generalization in streaming tasks where models are trained on short sequences but must operate over much longer, unknown horizons under bounded memory.
arXiv:2601. 22108v2 Announce Type: replace-cross Abstract: Continued pretraining is optimized with fixed self-supervised tasks but selected by downstream performance, creating a coarse feedback loop in which practitioners evaluate checkpoints, change data mixtures or objectives, and restart runs, while individual updates remain blind to target capabilities.
arXiv:2608. 11746v1 Announce Type: new Abstract: Modern systems are increasingly expected to transfer across tasks not specified during training.
AdaKerNet is a task‑adaptive neural kernel decoder that operates on frozen multimodal representations from large foundation models, without requiring access to the models’ parameters. It learns Lipschitz‑controlled multimodal features, a reference kernel providing a soft structural prior, and a lightweight nonlinear predictor that deforms this structure. Experiments on four multimodal large language models and diverse input modalities show consistent improvements over baseline decoders, achieving up to 41% error reduction in scarce‑label settings.
arXiv:2603. 15553v2 Announce Type: replace-cross Abstract: The landscape of self-supervised learning (SSL) is currently dominated by generative approaches (e.