arXiv AI

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents

arXiv:2607. 07985v1 Announce Type: cross Abstract: We report the empirical reliability of Gemini models as audio judges that score full-duplex agent conversations directly from the raw stereo waveform, tested across three models in the Gemini family: 2.

arXiv Machine Learning
Aug 28

Interpretable, Fairly Evaluated Automated L2 Speaking Assessment that Beats the Single-Human Ceiling and Why Pause Encoding Does Not Change LLM Fluency Scores

The paper presents an interpretable, fair, and accurately benchmarked automated system for assessing second‑language English speaking. Using a hybrid of feature‑based speech‑timing metrics and a large language model (LLM) fluency judgment, the system achieves a Spearman correlation of 0.818 with the ICNALE Global Rating Archive, outperforming 81 % of trained human raters. A controlled study shows that encoding pauses into the LLM prompt does not meaningfully affect fluency scores, indicating that the system’s fluency signal derives from measurable speech‑timing features.

By Eichi Uehara
arXiv Machine Learning
Sep 15

The Limits of Reference-Free Speech Quality Metrics as Evaluators and Rewards on Modern Text-to-Speech

arXiv:2609.13150v1 Announce Type: cross Abstract: Reference-free quality predictors such as UTMOS, DNSMOS and SCOREQ are the de facto automatic evaluators for text-to-speech (TTS) and are increasingl...

By Antonis Asonitis, Juan Pablo Zuluaga Gomez, Francesco Verdini, Aref Farhadipour, Marzieh Razavi, Pierre-Edouard Honnet, Vijeta Avijeet
arXiv AI
2d ago

Multi-Party Backchannel Prediction: a Diagnosis, a Benchmark, and a Ceiling

The paper introduces a multi‑party backchannel prediction benchmark built from the AMI meeting corpus, featuring 682 masked‑listener views, 190 speakers, and 18,697 backchannel events. A state‑of‑the‑art dyadic model performs at chance when applied zero‑shot to meetings, but its frozen acoustic features are still informative, and retraining improves performance to an AUROC of 0.751. The study reveals that listener conditioning helps only for listeners seen during training, that speaker identity is entangled with useful cues, and that backchannel rates vary significantly across individuals, prompting the authors to report both AUROC and event‑F1 metrics. whyItMatters":"The benchmark and evaluation tools provide a standardized, person‑disjoint testbed for advancing multi‑party backchannel prediction research."

By Mohammed Hafsati, Ahmed Loughzali
arXiv AI
6d ago

Audio LLMs Know When They Can't Hear You

The paper investigates whether audio large language models (Audio LLMs) can detect when their own transcriptions are unreliable. It finds that the models are poor at self-assessment and that existing methods offer limited detection. By leveraging audio-encoder representations, the authors develop a lightweight predictor that accurately flags unreliable transcriptions and can prompt user clarification without altering the underlying model.

By Amirhosein Javadi, Richa Dixit, Mehrdad Farajtabar, Minsik Cho, Devang Naik, Mohammad Samragh
Hugging Face Trending Papers
Aug 3

Can Foundation Models Hear What Made That Sound? A Tiered Benchmark of Audio-Language Models and Traditional Classifiers for Closed-Set Sound Source Identification

We benchmark eleven audio classification methods: five task-aware closed-set LLMs (four Gemini models plus open-weight Kimi-Audio-7B-Instruct), four fixed-vocabulary taggers (YAMNet, PANNs, Whisper-AT, and SSLAM), a zero-shot audio-text model (CLAP), and an audio-grounded LLM (BAT). We evaluate them on a closed-set sound-source identification task over 2,242 clips spanning 23 fine-grained classes and 11 categories.

arXiv Machine Learning
Sep 10

Clean Accuracy Does Not Guarantee Provenance Robustness: A Prospective Codec-Stress Evaluation of Audio Attribution

The study evaluates audio provenance attribution systems, showing that high clean‑benchmark accuracy does not translate to robustness after codec compression. Using a prospectively registered protocol, the authors measured closed‑set attribution performance on two corpora after single‑stage codec transport, finding significant degradation—up to 70.3 Macro‑F1 points for WavLM‑Base+ and 61.0 for W2V2‑BERT 2.0—depending on codec settings and representation. The results demonstrate that clean accuracy alone cannot guarantee deployment robustness across different codecs and representations.

By Gang Shi (Independent Researcher)