arXiv:2606. 25225v1 Announce Type: cross Abstract: Self-supervised learning from large-scale video data has emerged as a dominant paradigm for visual representation learning.
By Revant Teotia, Adrien Bardes, Michael Rabbat, Sumit Chopra, Matthew J. Muckley, Nicolas Ballas
arXiv:2607. 08545v1 Announce Type: cross Abstract: End-to-end neural audio models achieve high-fidelity compression and generation.
By Nicole Cosme-Clifford
The paper introduces a generalized multimodal foundation model that can handle arbitrary combinations of modalities and prediction tasks. It trains on large-scale synthetic multimodal datasets with diverse causal structures to learn transferable multimodal correlations. Experiments on 18 real-world datasets across 12 modalities and 11 tasks show competitive performance compared to specialized models without task-specific adaptation.
By Huizi Cui, Zongbo Han, Chenggong Ding, Naichuan Xiao, Jialong Yang, Jingdong Chen, Guangyu Wang, Qinghua Hu, Changqing Zhang
arXiv:2603. 07523v3 Announce Type: replace Abstract: Transferring knowledge by fine-tuning large-scale pre-trained networks has become a standard paradigm for downstream tasks, yet the knowledge of a pre-trained model is tightly coupled with monolithic architecture, which restricts flexible reuse across models of varying scales.
By Jianlu Shen, Fu Feng, Yucheng Xie, Jiaqi Lv, Xin Geng
AdaKerNet is a task‑adaptive neural kernel decoder that operates on frozen multimodal representations from large foundation models, without requiring access to the models’ parameters. It learns Lipschitz‑controlled multimodal features, a reference kernel providing a soft structural prior, and a lightweight nonlinear predictor that deforms this structure. Experiments on four multimodal large language models and diverse input modalities show consistent improvements over baseline decoders, achieving up to 41% error reduction in scarce‑label settings.
By Konstantinos D. Polyzos, Eleni Oikonomou, Tara Javidi
UniAE-MoE is a unified audio encoder that uses a Mixture-of-Experts architecture to model cross‑domain audio representations. It integrates encoder components from Qwen2‑Audio and Audio‑Flamingo 3, enhances them with SwiGLU and shared experts, and applies a two‑stage instruction‑tuning strategy along with task‑specific data scaling. The model achieves state‑of‑the‑art results on the XARES‑LLM benchmark (0.802) and tops the Interspeech 2026 Audio Encoder Capability Challenge, demonstrating strong generalization across speech, music, and general audio tasks.
By Shengbo Cai, Zhisheng Zhang, Zichao Nie, Jing Peng, Jingran Xie, Zhiyong Wu
RAST is a resolution‑aware transfer framework that addresses the mismatch between high‑resolution (HR) audio available during training and low‑resolution (LR) audio used at inference for human activity recognition. By compressing HR teacher representations while preserving token‑level information and neighborhood structure, RAST performs localized HR‑LR alignment. Experiments on the SAMoSA and AudioIMU datasets show that RAST consistently outperforms LR‑only training and direct teacher‑transfer baselines, improving LR‑only recognition by up to approximately 7.8% while requiring only LR audio at inference.
By Ji Hwan Park, Gautham Krishna Gudur, Yufei Shen, Dawei Liang, Edison Thomaz
The paper investigates whether the sparsity of Mixture-of-Experts (MoE) models leads to intrinsic semantic organization across modalities and domains. It shows that experts naturally specialize semantically even without explicit modular training. The authors propose ExpertLens, a data‑free method that decodes router weights to identify domain‑specialized experts, enabling selective fine‑tuning that matches or exceeds full fine‑tuning while updating only 21.7–47.0% of parameters and achieving a 4.0× speedup, outperforming LoRA in both performance and efficiency.
By Damiano Marsili, Raphi Kang, Aditya Mehta, Pietro Perona, Georgia Gkioxari
arXiv:2503. 06211v3 Announce Type: replace-cross Abstract: Text-pretrained language models (LMs) encode rich world knowledge, but adapting them to process and generate perceptual modalities such as audio and images while effectively leveraging that knowledge remains challenging.
By Santiago Cuervo, Adel Moumen, Yanis Labrak, Sameer Khurana, Antoine Laurent, Mickael Rouvier, Phil Woodland, Ricard Marxer
arXiv:2608. 14819v1 Announce Type: cross Abstract: Music foundation models are commonly used as frozen audio feature extractors, yet selecting which layer to extract from remains largely heuristic.
By Angelos-Nikolaos Kanatas, Yuexuan Kong, Pablo Alonso-Jim\'enez, Xavier Serra, Dmitry Bogdanov
CIG-MAE is a self‑supervised framework for WiFi‑based human action recognition that uses a cross‑modal masked autoencoder to reconstruct both amplitude and phase of Channel State Information. It introduces an adaptive, information‑guided masking strategy that focuses on high‑density time‑frequency regions and employs a Barlow Twins regularizer to align cross‑modal representations without negative samples. Experiments on three public datasets show that CIG‑MAE outperforms state‑of‑the‑art SSL methods and even surpasses a fully supervised baseline, highlighting its data efficiency, robustness, and generalization.
By Gang Liu, Yanling Hao, Yixuan Zou
AudioFuse is a hybrid architecture that jointly learns from spectrograms and raw waveforms to classify phonocardiograms. It combines a wide-and-shallow Vision Transformer for spectral features with a shallow 1D CNN for temporal waveforms, reducing overfitting while capturing complementary information. On the PhysioNet 2016 dataset, AudioFuse achieves a state‑of‑the‑art ROC‑AUC of 0.8608 and shows superior robustness to domain shift on the PASCAL dataset, outperforming both spectrogram‑only and waveform‑only baselines.
By Md. Saiful Bari Siddiqui, Utsab Saha