arXiv:2608. 01481v1 Announce Type: new Abstract: Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2.
By Ilia Semenkov, Daria Kleeva, Ivan Dakhtin, Zarina Maksudova, Alex Ossadtchi
arXiv:2607. 11801v1 Announce Type: cross Abstract: Large audio-language models (LALMs) often underperform on fine-grained, non-semantic attributes of speech, such as a speaker's emotion, despite strong performance on speech content.
By Yu-Han Huang, Chih-Kai Yang, Ke-Han Lu, An-Yu Cheng, Hung-yi Lee
Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2. 0 audio embeddings.
arXiv:2607. 22413v1 Announce Type: cross Abstract: Sample retrieval tools can help composers find harmonically compatible material, but querying from a fixed reference sample becomes less informative as arrangements evolve and the harmonic context shifts with each musical decision.
By Austin Rockman
arXiv:2511. 05350v3 Announce Type: replace-cross Abstract: We argue that training autoencoders to reconstruct inputs from noised versions of their encodings, when combined with perceptually motivated losses, yields encodings that are structured according to a perceptual hierarchy.
By Mathias Rose Bjare, Giorgia Cantisani, Marco Pasini, Stefan Lattner, Gerhard Widmer
arXiv:2601. 18904v3 Announce Type: replace-cross Abstract: Generative AI for speech and audio is increasingly expected to serve users across languages, cultures, and communities, yet current auditory Large Language Models (LLMs) are still largely trained and evaluated on high-resource data.
By Haolong Zheng, Siyin Wang, Zengrui Jin, Mark Hasegawa-Johnson
arXiv:2606. 17301v1 Announce Type: cross Abstract: Search, a foundational operation in computer science, maps a query to a matching item in a collection.
By Muhammad Taimoor Haseeb, Ahmad Hammoudeh, Gus Xia
The paper introduces AIB, the first benchmark of auditory perceptual illusion tasks designed for Large Audio Language Models (LALMs). It covers ten representative music, sound, and speech‑based illusory phenomena, each annotated for knowledge‑based priors, and evaluates models alongside controlled human listening studies. Results reveal that while LALMs are generally signal‑faithful on low‑level acoustic cues, some models show more human‑like responses when linguistic or musical priors are involved, yet none fully match human perceptual patterns.
By Hayoon Kim, Eunice Hong, Kyogu Lee
arXiv:2609.36577v1 Announce Type: cross
Abstract: Audio large language models (ALLMs) can reason about the content of audio recordings to perform complex tasks. However, these capabilities usually co...
By Zhenhong Zhou, Xuanyue Zhao, Youji Liu, Yuanhe Zhang, Xiaoyu Ma, Lianyu Hu, Yang Liu
AnchorPrompt is an adaptation technique for large audio‑language models that keeps the base model frozen and learns a single block of prompt vectors inserted at the decoder input. By training these prompts through self‑distillation on diverse audio and text perturbations, the method improves answer consistency and reduces hallucinations across multiple benchmarks. The approach is perturbation‑agnostic at inference, enabling zero‑shot transfer to unseen distortions such as reverberation and choice permutations.
By Pooneh Mousavi, Amir Ivry, Mirco Ravanelli, Cem Subakan
arXiv:2607. 08545v1 Announce Type: cross Abstract: End-to-end neural audio models achieve high-fidelity compression and generation.
By Nicole Cosme-Clifford
arXiv:2607. 00247v1 Announce Type: cross Abstract: Large audio-language models (LALMs) frequently hallucinate by overriding acoustic evidence with language priors.
By Aaron Isidore Grace, Zhouyuan Huo, Weiran Wang