arXiv AI

Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models

arXiv:2607. 11801v1 Announce Type: cross Abstract: Large audio-language models (LALMs) often underperform on fine-grained, non-semantic attributes of speech, such as a speaker's emotion, despite strong performance on speech content.

arXiv AI
Aug 10

MetaSICL: Globalizing Auditory LLMs for Underserved Speakers and Languages via Meta Speech In-Context Learning

arXiv:2601. 18904v3 Announce Type: replace-cross Abstract: Generative AI for speech and audio is increasingly expected to serve users across languages, cultures, and communities, yet current auditory Large Language Models (LLMs) are still largely trained and evaluated on high-resource data.

By Haolong Zheng, Siyin Wang, Zengrui Jin, Mark Hasegawa-Johnson
arXiv AI
Sep 2

Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models

The paper investigates whether audio‑language models capture paralinguistic cues beyond spoken content. Using the Expresso dataset and four open‑source models, the authors trace how speaking style information is encoded in the late layers of the audio encoder but is degraded before reaching the final output. They find that some models are content‑driven while others are acoustic‑driven, revealing a gap between what is encoded and what is utilized in current audio‑language models.

By Bhuvan Koduru, Dareen Safar B Alharthi, Rita Singh, Bhiksha Raj
arXiv AI
Sep 7

Tracing Audio Grounding and Answer Selection in Audio LLMs

The paper investigates how Audio Large Language Models (Audio LLMs) actually use audio input to determine answers, rather than relying on textual cues. It finds that replacing audio with silence or unrelated audio degrades performance more after training than before, that acoustic information shapes representations in early-to-middle layers and influences final predictions in middle-to-late layers, and that training impacts specific layer bands most strongly. These observations offer a mechanistic view of how training enhances the use of acoustic evidence in Audio LLMs.

By Hyebin Cho, Suho Yoo, Jihoo Jung, Joon Son Chung
arXiv Machine Learning
1d ago

AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models

AnchorPrompt is an adaptation technique for large audio‑language models that keeps the base model frozen and learns a single block of prompt vectors inserted at the decoder input. By training these prompts through self‑distillation on diverse audio and text perturbations, the method improves answer consistency and reduces hallucinations across multiple benchmarks. The approach is perturbation‑agnostic at inference, enabling zero‑shot transfer to unseen distortions such as reverberation and choice permutations.

By Pooneh Mousavi, Amir Ivry, Mirco Ravanelli, Cem Subakan