arXiv AI

Multi-agent Auditory Scene Analysis: Improved Localization Speed and Robustness by Multi-beamformed Speech Quality Feedback

arXiv Computation and Language
Sep 10

EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents

arXiv:2605.13841v3 Announce Type: replace-cross Abstract: Voice agents are increasingly deployed across enterprise applications. However, no existing benchmark jointly addresses realistic conversatio...

By Tara Bogavelli, Gabrielle Gauthier Melan\c{c}on, Katrina Stankiewicz, Oluwanifemi Bamgbose, Fanny Riols, Hoang H. Nguyen, Raghav Mehndiratta, Lindsay Devon Brin, Joseph Marinier, Hari Subramani, Anil Madamala, Sridhar Krishna Nemala, Srinivas Sunkara
arXiv AI
Sep 1

KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs

The paper introduces KVoiceBench, KOpenAudioBench, and KMMAU—three Korean speech benchmarks created through agent-driven frameworks that adapt existing SpokenQA and ASR resources into Korean SpokenQA and audio understanding tasks. These benchmarks total 12,345 samples and are publicly released to evaluate SpeechLMs beyond English. The authors benchmark eight recent SpeechLMs, revealing significant English‑Korean performance gaps and divergent rankings between SpokenQA and audio understanding, highlighting multilingual weaknesses not apparent in English-only tests.

By Haechan Kim, Seungjun Chung, Inkyu Park, Jihoo Lee, Jonghyun Lee
arXiv AI
6d ago

Evaluating Real-Time Voice Agents: From Component Quality to Grounded Outcomes

The paper reviews the fragmented literature on real‑time voice agents, noting that architecture, turn‑taking, and agentic evaluation communities rarely cite each other. It presents three evidence‑based claims: (1) architecture choice is a deployment constraint rather than a definitive solution, (2) evaluation has shifted from component quality to grounded outcomes, and (3) the dyadic assumption is breaking down as multiparty turn‑taking and reasoning become essential. The authors propose the TRG reporting standard to characterize agents by timing, recovery, and state‑verified outcomes, with an optional fourth axis for multiparty contexts.

By Shivam Negi, Arpit Rawat, Rashi Jain
arXiv AI
Sep 7

Scalable Context Orchestration for Serving LLMs Over Voice

The paper introduces llmovoice, a middleware that explicitly models voice context for large language model (LLM) serving in voice AI applications. By incorporating speaking rate, background noise, packet loss, and other paralinguistic factors into a bounded context, llmovoice guides the LLM to generate more aligned responses. Experiments show significant reductions in speaking‑rate errors, false interruptions, and model usage costs, especially in long voice sessions.

By Linyi Jiang, Silvery D. Fu, Yifei Zhu
arXiv AI
6d ago

Inquesto Score: A reliability Protocol For Voice Agents

The paper introduces Inquesto Score (IS), a protocol that measures voice‑agent reliability by calculating the percentage of calls that reach the caller’s goal without functional failure. IS defines explicit failure events and severity levels, evaluates timing, semantic, and state‑dependent failures using audio, scenario predicates, tool traces, and a pinned open‑model judge, and provides diagnostic views on behavior, acoustic robustness, identity handling, and speaker groups. The authors evaluate IS v0.1 on 30 scenarios, three acoustic conditions, four speaker groups, and 306 calls per agent across 13 configurations, demonstrating that reliable measurement requires evidence beyond transcripts and explicit treatment of deployment conditions.

By Massa Baali, Bhiksha Raj