arXiv Computation and Language

CHiME-9 ECHI: A Machine Learning Challenge for Enhancing Conversations to Address Hearing Impairment

arXiv Computation and Language
Sep 23

Challenges of Multi-Speaker Extraction for Real Conversational Speech Enhancement

The paper addresses challenges in extracting target and multiple speakers from real conversational speech, noting that real conversations contain more silence and enrolment samples that differ from the target speech. It introduces a new loss function that reduces the impact of excess silence during training, yielding improvements in STOI (from 0.55 to 0.60) and frequency‑weighted segmental SNR (from 4.35 to 5.12). The study also investigates how mismatches between enrolment and target speech affect performance.

By Robert Sutherland, Stefan Goetze, Jon Barker
arXiv AI
Jun 4

The Differentiable Auditory Loop (DAL): An ML Framework for Hyper-Personalized Hearing Aids

arXiv:2606. 04103v1 Announce Type: cross Abstract: Conventional hearing aids rely on fixed, frequency-dependent amplification and compression to manage reduced sensitivity, which often fails to provide sufficient listening support in complex environments, such as situations with multiple speakers (the ``cocktail party'' problem).

By Alejandro Ballesta Rosen, Jason Mikiel-Hunter, Julian Maclaren, Jack Collins, Richard F. Lyon, Simon Carlile
arXiv AI
Jul 17

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems

arXiv:2607. 14846v1 Announce Type: cross Abstract: Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual representation.

By David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer, Jakub Piotr C{\l}apa, Theo Lebryk, Jens Madsen, Olya Ossipova, Sharath Rao, Hoon Shin, Tigran Soghbatyan, Georg Streich, Rashish Tandon, Panagiotis Tzirakis
arXiv Computation and Language
Sep 10

EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents

arXiv:2605.13841v3 Announce Type: replace-cross Abstract: Voice agents are increasingly deployed across enterprise applications. However, no existing benchmark jointly addresses realistic conversatio...

By Tara Bogavelli, Gabrielle Gauthier Melan\c{c}on, Katrina Stankiewicz, Oluwanifemi Bamgbose, Fanny Riols, Hoang H. Nguyen, Raghav Mehndiratta, Lindsay Devon Brin, Joseph Marinier, Hari Subramani, Anil Madamala, Sridhar Krishna Nemala, Srinivas Sunkara
arXiv Machine Learning
Sep 15

The Limits of Reference-Free Speech Quality Metrics as Evaluators and Rewards on Modern Text-to-Speech

arXiv:2609.13150v1 Announce Type: cross Abstract: Reference-free quality predictors such as UTMOS, DNSMOS and SCOREQ are the de facto automatic evaluators for text-to-speech (TTS) and are increasingl...

By Antonis Asonitis, Juan Pablo Zuluaga Gomez, Francesco Verdini, Aref Farhadipour, Marzieh Razavi, Pierre-Edouard Honnet, Vijeta Avijeet
arXiv AI
6d ago

Inquesto Score: A reliability Protocol For Voice Agents

The paper introduces Inquesto Score (IS), a protocol that measures voice‑agent reliability by calculating the percentage of calls that reach the caller’s goal without functional failure. IS defines explicit failure events and severity levels, evaluates timing, semantic, and state‑dependent failures using audio, scenario predicates, tool traces, and a pinned open‑model judge, and provides diagnostic views on behavior, acoustic robustness, identity handling, and speaker groups. The authors evaluate IS v0.1 on 30 scenarios, three acoustic conditions, four speaker groups, and 306 calls per agent across 13 configurations, demonstrating that reliable measurement requires evidence beyond transcripts and explicit treatment of deployment conditions.

By Massa Baali, Bhiksha Raj
arXiv AI
2d ago

Multi-Party Backchannel Prediction: a Diagnosis, a Benchmark, and a Ceiling

The paper introduces a multi‑party backchannel prediction benchmark built from the AMI meeting corpus, featuring 682 masked‑listener views, 190 speakers, and 18,697 backchannel events. A state‑of‑the‑art dyadic model performs at chance when applied zero‑shot to meetings, but its frozen acoustic features are still informative, and retraining improves performance to an AUROC of 0.751. The study reveals that listener conditioning helps only for listeners seen during training, that speaker identity is entangled with useful cues, and that backchannel rates vary significantly across individuals, prompting the authors to report both AUROC and event‑F1 metrics. whyItMatters":"The benchmark and evaluation tools provide a standardized, person‑disjoint testbed for advancing multi‑party backchannel prediction research."

By Mohammed Hafsati, Ahmed Loughzali