arXiv AI

RAG-Audio: Retrieval-Augmented Generation for Faithful Brain-to-Audio Reconstruction

arXiv:2608. 09331v1 Announce Type: cross Abstract: Brain-to-audio reconstruction is limited by \emph{prior domination}: when a pretrained generator is conditioned on a weak neural signal, it produces realistic but stimulus-inaccurate audio.

arXiv AI
Aug 10

MetaSICL: Globalizing Auditory LLMs for Underserved Speakers and Languages via Meta Speech In-Context Learning

arXiv:2601. 18904v3 Announce Type: replace-cross Abstract: Generative AI for speech and audio is increasingly expected to serve users across languages, cultures, and communities, yet current auditory Large Language Models (LLMs) are still largely trained and evaluated on high-resource data.

By Haolong Zheng, Siyin Wang, Zengrui Jin, Mark Hasegawa-Johnson
arXiv AI
Sep 3

Auditory Illusion Benchmark for Large Audio Language Models

The paper introduces AIB, the first benchmark of auditory perceptual illusion tasks designed for Large Audio Language Models (LALMs). It covers ten representative music, sound, and speech‑based illusory phenomena, each annotated for knowledge‑based priors, and evaluates models alongside controlled human listening studies. Results reveal that while LALMs are generally signal‑faithful on low‑level acoustic cues, some models show more human‑like responses when linguistic or musical priors are involved, yet none fully match human perceptual patterns.

By Hayoon Kim, Eunice Hong, Kyogu Lee
arXiv Machine Learning
1d ago

AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models

AnchorPrompt is an adaptation technique for large audio‑language models that keeps the base model frozen and learns a single block of prompt vectors inserted at the decoder input. By training these prompts through self‑distillation on diverse audio and text perturbations, the method improves answer consistency and reduces hallucinations across multiple benchmarks. The approach is perturbation‑agnostic at inference, enabling zero‑shot transfer to unseen distortions such as reverberation and choice permutations.

By Pooneh Mousavi, Amir Ivry, Mirco Ravanelli, Cem Subakan