arXiv AI By Massa Baali, Bhiksha Raj

Inquesto Score: A reliability Protocol For Voice Agents

Read the original on arXiv AI →

The paper introduces Inquesto Score (IS), a protocol that measures voice‑agent reliability by calculating the percentage of calls that reach the caller’s goal without functional failure. IS defines explicit failure events and severity levels, evaluates timing, semantic, and state‑dependent failures using audio, scenario predicates, tool traces, and a pinned open‑model judge, and provides diagnostic views on behavior, acoustic robustness, identity handling, and speaker groups. The authors evaluate IS v0.1 on 30 scenarios, three acoustic conditions, four speaker groups, and 306 calls per agent across 13 configurations, demonstrating that reliable measurement requires evidence beyond transcripts and explicit treatment of deployment conditions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 10

EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents

arXiv:2605.13841v3 Announce Type: replace-cross Abstract: Voice agents are increasingly deployed across enterprise applications. However, no existing benchmark jointly addresses realistic conversatio...

By Tara Bogavelli, Gabrielle Gauthier Melan\c{c}on, Katrina Stankiewicz, Oluwanifemi Bamgbose, Fanny Riols, Hoang H. Nguyen, Raghav Mehndiratta, Lindsay Devon Brin, Joseph Marinier, Hari Subramani, Anil Madamala, Sridhar Krishna Nemala, Srinivas Sunkara
arXiv AI
Aug 26

Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight

The paper evaluates the use of large language models (LLMs) as judges for assessing conversational voice agents, comparing human judgments with GPT‑4.1 and GPT‑5 across telecom and retail interactions. It examines agreement, metric‑level correlations, and consistency across three evaluation configurations (p0, p1, p2) to determine how reliably LLMs can judge conversational quality and safety. The study finds that LLM‑based evaluation can be effective but its reliability varies by metric and configuration, suggesting a hybrid approach where LLMs handle scalable assessment while humans focus on metrics requiring contextual interpretation.

By Anupam Purwar, Shashank Singh, Kritika Srivastava
arXiv AI
4d ago

Evaluating Real-Time Voice Agents: From Component Quality to Grounded Outcomes

The paper reviews the fragmented literature on real‑time voice agents, noting that architecture, turn‑taking, and agentic evaluation communities rarely cite each other. It presents three evidence‑based claims: (1) architecture choice is a deployment constraint rather than a definitive solution, (2) evaluation has shifted from component quality to grounded outcomes, and (3) the dyadic assumption is breaking down as multiparty turn‑taking and reasoning become essential. The authors propose the TRG reporting standard to characterize agents by timing, recovery, and state‑verified outcomes, with an optional fourth axis for multiparty contexts.

By Shivam Negi, Arpit Rawat, Rashi Jain
arXiv Computation and Language
Sep 7

Auditing Bias and Safety in Voice AI Customer Care

The paper introduces a validation‑gated audit framework for voice AI customer‑care systems, treating them as stateful, multi‑turn, tool‑mediated interactions where bias and safety can manifest as added burdens before a final decision. The framework distinguishes between native speech‑to‑speech, cascaded ASR‑to‑LM‑to‑TTS, and hybrid architectures, and applies matched service facts across controlled caller presentation conditions to validate fact invariance, presentation cues, artifacts, and acoustic measurements. It outlines seven validation gates, a six‑family metric set, and demonstrates the approach with a synthetic refund‑dispute audit example, while noting that production results are withheld until the protocol is satisfied.

By Vignesh Ethiraj, Ashwath David