Structural Bottlenecks on Frequency Representation in End-to-End Audio Models
arXiv:2607. 08545v1 Announce Type: cross Abstract: End-to-end neural audio models achieve high-fidelity compression and generation.
arXiv:2607. 16027v1 Announce Type: new Abstract: Introduction: Biological systems face anatomical and metabolic constraints, including costly synaptic maintenance and limited connectivity.
arXiv:2607. 08545v1 Announce Type: cross Abstract: End-to-end neural audio models achieve high-fidelity compression and generation.
arXiv:2608. 09227v1 Announce Type: new Abstract: Omnimodal language models (OLMs) enable unified audio-visual understanding, but processing long joint token sequences makes inference computationally prohibitive.
AdaKerNet is a task‑adaptive neural kernel decoder that operates on frozen multimodal representations from large foundation models, without requiring access to the models’ parameters. It learns Lipschitz‑controlled multimodal features, a reference kernel providing a soft structural prior, and a lightweight nonlinear predictor that deforms this structure. Experiments on four multimodal large language models and diverse input modalities show consistent improvements over baseline decoders, achieving up to 41% error reduction in scarce‑label settings.
arXiv:2606. 22790v2 Announce Type: replace-cross Abstract: In this paper, we investigate the tradeoffs between compute allocation and model performance for two speech processing tasks: Automatic Speech Recognition (ASR) and Speech Emotion Recognition (SER).
arXiv:2608. 01481v1 Announce Type: new Abstract: Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2.
Omnimodal language models (OLMs) enable unified audio-visual understanding, but processing long joint token sequences makes inference computationally prohibitive. While recent token compression methods attempt to alleviate this burden, compressing modalities in isolation often destroys the temporal cross-modal anchors necessary for coherent reasoning.
arXiv:2606. 16408v1 Announce Type: new Abstract: We introduce MUNI, an end-to-end multimodal latent diffusion framework for any-to-any generation that unifies subset-conditioned cross-modal generation and unconditional joint sampling through a shared stochastic latent.
The paper introduces Supervised Spike Agreement-Dependent Plasticity (Supervised SADP), a gradient‑free Hebbian learning rule that embeds class labels directly into spike‑driven plasticity. SADP trains output neurons with a supervised Hebbian rule and hidden neurons by measuring Cohen’s kappa agreement with the correct‑class output spike train, optionally aggregating over temporal offsets (K‑shift). Across six benchmark and medical imaging datasets, SADP consistently outperforms reward‑modulated STDP, achieving higher accuracy (e.g., 86.46 % on MNIST) and faster training (up to 2.86× speedup).
arXiv:2606. 09853v1 Announce Type: new Abstract: A central objective in multimodal learning is to capture synergy: task-relevant information that arises only from the joint use of multiple modalities, and is not available from any single modality alone.
arXiv:2608. 08569v1 Announce Type: new Abstract: Recent advancements in Speech Large Language Models have demonstrated remarkable capabilities in understanding complex audio tasks.
arXiv:2509. 24039v2 Announce Type: replace-cross Abstract: If topography is a fundamental feature of the brain, it should influence both how neurons are arranged in space (i.
The paper introduces modality‑gated deep adapters, a parameter‑efficient method for adding new modalities to a frozen multimodal embedding language model without altering its existing outputs. These adapters are bottleneck modules attached to each decoder layer, grouped into modality‑specific packs that activate only during encoding of their own modality, ensuring exact preservation of the base model’s computation graph. Experiments on a 2B base model show significant gains in audio‑to‑text and thermal‑to‑text retrieval metrics, and the authors release the audio and thermal packs along with training and evaluation code.