arXiv AI

Audio LLMs Know When They Can't Hear You

The paper investigates whether audio large language models (Audio LLMs) can detect when their own transcriptions are unreliable. It finds that the models are poor at self-assessment and that existing methods offer limited detection. By leveraging audio-encoder representations, the authors develop a lightweight predictor that accurately flags unreliable transcriptions and can prompt user clarification without altering the underlying model.

arXiv Computation and Language
6d ago

Proactive for Uncertainty: Cause-Aware Error Diagnosis and Interactive Clarification for Spoken Dialogue Systems

The paper introduces a cause-aware error recovery framework for cascaded Automatic Speech Recognition – Large Language Model (ASR‑LLM) pipelines in Spoken Dialogue Systems. It replaces simple ASR confidence filtering with precision‑focused detectors that use deep ASR latent representations to classify token‑level errors into perception, comprehension, and deletion failures. This fine‑grained diagnosis enables the LLM to execute targeted, multi‑turn clarification strategies, leading to a more than two‑fold increase in recall on domain‑shift errors and significant reductions in word error rate and downstream task errors across varied accents, distortions, and domains.

By Yizhou Peng, Ziyang Ma, Changsong Liu, Yi-Wen Chao, Xie Chen, Eng Siong Chng
arXiv AI
Sep 2

Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models

The paper investigates whether audio‑language models capture paralinguistic cues beyond spoken content. Using the Expresso dataset and four open‑source models, the authors trace how speaking style information is encoded in the late layers of the audio encoder but is degraded before reaching the final output. They find that some models are content‑driven while others are acoustic‑driven, revealing a gap between what is encoded and what is utilized in current audio‑language models.

By Bhuvan Koduru, Dareen Safar B Alharthi, Rita Singh, Bhiksha Raj
arXiv Computation and Language
Aug 28

FireRedAudio: A General-Purpose Audio Language Model with Decoupled Continuous Representations for Understanding and Generation

FireRedAudio is a 9‑billion‑parameter audio language model that separates continuous input representations for audio understanding and speech generation, enabling a single autoregressive LLM to perform tasks such as ASR, zero‑shot TTS, Instruct TTS, and semantic/acoustic speech editing. The model uses a dedicated Audio Encoder for recognition and a RedAE‑based pathway for generation, with the LLM directly generating text or conditioning a flow‑matching DiT to produce acoustic latents. Evaluations show competitive or leading performance in multilingual ASR, content‑accurate zero‑shot TTS, strong instruction following, and significant improvements in speech editing over prior work.

By Feiyu Shen, Fenglong Xie, Junjie Li, Kun Xie, Lei Xie, Xu Tang, Xuelong Geng, Yan Jia, Yao Hu, Yichen Han, Yichen Wu, Ziqi Dai, Junjie Chen, Kai Huang, Manzhen Wei, Yixuan Li
arXiv Computation and Language
3d ago

Asymmetric Classifier-Free Guidance for Target-Speaker ASR

The paper introduces asymmetric classifier‑free guidance (CFG) for target‑speaker ASR using Whisper, where a speaker‑conditioned branch predicts the target transcript and a speaker‑unconditioned branch predicts serialized multi‑speaker transcripts. CFG modulates the influence of speaker conditioning during decoding via a single guidance scale, which is first set globally on development data and then refined per utterance by a lightweight encoder‑based predictor while keeping the recognition model fixed. The resulting system yields up to 21.8% relative WER reduction over a condition‑only baseline and 5.6% over standard conditional decoding under domain shifts.

By Yiwen Guan, Jacob Whitehill
arXiv Machine Learning
Sep 15

The Limits of Reference-Free Speech Quality Metrics as Evaluators and Rewards on Modern Text-to-Speech

arXiv:2609.13150v1 Announce Type: cross Abstract: Reference-free quality predictors such as UTMOS, DNSMOS and SCOREQ are the de facto automatic evaluators for text-to-speech (TTS) and are increasingl...

By Antonis Asonitis, Juan Pablo Zuluaga Gomez, Francesco Verdini, Aref Farhadipour, Marzieh Razavi, Pierre-Edouard Honnet, Vijeta Avijeet
arXiv Machine Learning
Sep 2

MRMAD: A Multi-Round Multi-Audio Benchmark for Evaluating Acoustic Degradation Perception in Large Audio-Language Models

arXiv:2608.22236v2 Announce Type: replace-cross Abstract: Large audio-language models (LALMs) have shown promising progress in understanding speech, music, and general sound events, yet their ability...

By Yize Li, Ningyuan Yang, Sile Yin, Sindhuja Thogarrati, Sung-En Chang, Andrew C. Singer, Xue Lin, Chuan-Che Huang, Shuo Zhang
arXiv AI
Jun 18

Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation

arXiv:2603. 10827v2 Announce Type: replace-cross Abstract: Speech-aware large language models (LLMs) can accept speech inputs, yet their training objectives largely emphasize linguistic content or specific fields such as emotions or the speaker's gender, leaving it unclear whether they encode speaker identity.

By Thomas Thebaud, Yuzhe Wang, Laureano Moro-Velazquez, Jesus Villalba-Lopez, Najim Dehak