arXiv AI

Inquesto Score: A reliability Protocol For Voice Agents

The paper introduces Inquesto Score (IS), a protocol that measures voice‑agent reliability by calculating the percentage of calls that reach the caller’s goal without functional failure. IS defines explicit failure events and severity levels, evaluates timing, semantic, and state‑dependent failures using audio, scenario predicates, tool traces, and a pinned open‑model judge, and provides diagnostic views on behavior, acoustic robustness, identity handling, and speaker groups. The authors evaluate IS v0.1 on 30 scenarios, three acoustic conditions, four speaker groups, and 306 calls per agent across 13 configurations, demonstrating that reliable measurement requires evidence beyond transcripts and explicit treatment of deployment conditions.

arXiv Computation and Language
Sep 10

EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents

arXiv:2605.13841v3 Announce Type: replace-cross Abstract: Voice agents are increasingly deployed across enterprise applications. However, no existing benchmark jointly addresses realistic conversatio...

By Tara Bogavelli, Gabrielle Gauthier Melan\c{c}on, Katrina Stankiewicz, Oluwanifemi Bamgbose, Fanny Riols, Hoang H. Nguyen, Raghav Mehndiratta, Lindsay Devon Brin, Joseph Marinier, Hari Subramani, Anil Madamala, Sridhar Krishna Nemala, Srinivas Sunkara
arXiv AI
Aug 26

Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight

The paper evaluates the use of large language models (LLMs) as judges for assessing conversational voice agents, comparing human judgments with GPT‑4.1 and GPT‑5 across telecom and retail interactions. It examines agreement, metric‑level correlations, and consistency across three evaluation configurations (p0, p1, p2) to determine how reliably LLMs can judge conversational quality and safety. The study finds that LLM‑based evaluation can be effective but its reliability varies by metric and configuration, suggesting a hybrid approach where LLMs handle scalable assessment while humans focus on metrics requiring contextual interpretation.

By Anupam Purwar, Shashank Singh, Kritika Srivastava
arXiv AI
4d ago

Evaluating Real-Time Voice Agents: From Component Quality to Grounded Outcomes

The paper reviews the fragmented literature on real‑time voice agents, noting that architecture, turn‑taking, and agentic evaluation communities rarely cite each other. It presents three evidence‑based claims: (1) architecture choice is a deployment constraint rather than a definitive solution, (2) evaluation has shifted from component quality to grounded outcomes, and (3) the dyadic assumption is breaking down as multiparty turn‑taking and reasoning become essential. The authors propose the TRG reporting standard to characterize agents by timing, recovery, and state‑verified outcomes, with an optional fourth axis for multiparty contexts.

By Shivam Negi, Arpit Rawat, Rashi Jain
arXiv Computation and Language
Sep 7

Auditing Bias and Safety in Voice AI Customer Care

The paper introduces a validation‑gated audit framework for voice AI customer‑care systems, treating them as stateful, multi‑turn, tool‑mediated interactions where bias and safety can manifest as added burdens before a final decision. The framework distinguishes between native speech‑to‑speech, cascaded ASR‑to‑LM‑to‑TTS, and hybrid architectures, and applies matched service facts across controlled caller presentation conditions to validate fact invariance, presentation cues, artifacts, and acoustic measurements. It outlines seven validation gates, a six‑family metric set, and demonstrates the approach with a synthetic refund‑dispute audit example, while noting that production results are withheld until the protocol is satisfied.

By Vignesh Ethiraj, Ashwath David
arXiv AI
Jul 17

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems

arXiv:2607. 14846v1 Announce Type: cross Abstract: Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual representation.

By David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer, Jakub Piotr C{\l}apa, Theo Lebryk, Jens Madsen, Olya Ossipova, Sharath Rao, Hoon Shin, Tigran Soghbatyan, Georg Streich, Rashish Tandon, Panagiotis Tzirakis
arXiv AI
Sep 7

Scalable Context Orchestration for Serving LLMs Over Voice

The paper introduces llmovoice, a middleware that explicitly models voice context for large language model (LLM) serving in voice AI applications. By incorporating speaking rate, background noise, packet loss, and other paralinguistic factors into a bounded context, llmovoice guides the LLM to generate more aligned responses. Experiments show significant reductions in speaking‑rate errors, false interruptions, and model usage costs, especially in long voice sessions.

By Linyi Jiang, Silvery D. Fu, Yifei Zhu