The paper evaluates the use of large language models (LLMs) as judges for assessing conversational voice agents, comparing human judgments with GPT‑4.1 and GPT‑5 across telecom and retail interactions. It examines agreement, metric‑level correlations, and consistency across three evaluation configurations (p0, p1, p2) to determine how reliably LLMs can judge conversational quality and safety. The study finds that LLM‑based evaluation can be effective but its reliability varies by metric and configuration, suggesting a hybrid approach where LLMs handle scalable assessment while humans focus on metrics requiring contextual interpretation.
By Anupam Purwar, Shashank Singh, Kritika Srivastava
arXiv:2606. 19595v1 Announce Type: cross Abstract: Voice agents deployed in structured workflows (customer service, healthcare scheduling, account management) must handle frequent user interruptions while maintaining progress through multi-step procedures.
By Ahmad Salimi, Wentao Ma, Yuzhi Tang, Dongming Shen, Mu Li, Alex Smola
The paper introduces Inquesto Score (IS), a protocol that measures voice‑agent reliability by calculating the percentage of calls that reach the caller’s goal without functional failure. IS defines explicit failure events and severity levels, evaluates timing, semantic, and state‑dependent failures using audio, scenario predicates, tool traces, and a pinned open‑model judge, and provides diagnostic views on behavior, acoustic robustness, identity handling, and speaker groups. The authors evaluate IS v0.1 on 30 scenarios, three acoustic conditions, four speaker groups, and 306 calls per agent across 13 configurations, demonstrating that reliable measurement requires evidence beyond transcripts and explicit treatment of deployment conditions.
By Massa Baali, Bhiksha Raj
arXiv:2609.24812v1 Announce Type: new
Abstract: Voice provides a natural and immediate interface for AI agents. Many settings in which voice agents could be useful, including meetings, households, an...
By Chenxu Xiong, Dongming Shen, Yuzhi Tang, Wentao Ma, Mu Li, Alex Smola
arXiv:2609.13076v1 Announce Type: cross
Abstract: Conversational voice agents have advanced significantly, offering increasingly natural human-machine interactions through both cascaded and end-to-en...
By Yi-Jen Shih, Shih-Yun Shan Kuan, Guan-Ting Lin, Kai-Wei Chang, Siddhant Arora, Shu-wen Yang, Abdelrahman Mohamed, Shinji Watanabe, Hung-yi Lee, David Harwath
arXiv:2609.38867v1 Announce Type: new
Abstract: Large language model (LLM) computer-use agents are typically evaluated with clean written instructions, despite speech being an increasingly popular in...
By Terumi Chiba, Guangzhi Sun, Zheqi Yuan, Chao Zhang
arXiv:2608. 10716v1 Announce Type: cross Abstract: Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consumers owing to the ease of the conversational modality over text.
By Aryan Vijay Bhosale, Harshit Rajgarhia, Akhil Pothanapalli, Asif Shaik, Abhishek Mukherji, Dinesh Manocha
DuplexSpeechBench-IFEval (DSB-IFEval) is a new benchmark that evaluates how full‑duplex voice agents follow implicit instructions during real‑time spoken interaction. It contains 1,038 test cases across eight assistant roles and tests five conditioning protocols, measuring floor management with an Instruction Adherence Score (IAS) and persona consistency with a Persona Adherence Score (PAS). Experiments on six speech systems reveal architecture‑dependent trade‑offs, showing that some models are more sensitive to explicit versus persona‑only instructions and that even when following conflicting directives, they struggle to override them under safety conflict.
By Puneet Mathur, Dinesh Manocha
MTVA-Bench is a new benchmark designed to evaluate the language model component of cascaded voice agents under realistic conditions. It simulates callers with an LLM, mocks backend tool responses, and scores both tool‑call correctness and conversational quality using two LLM judges, covering 49 agents, 490 scenarios, and 7 languages. The benchmark reveals that while models perform similarly on tool selection, they differ widely in argument handling, rule compliance, and dialogue quality, highlighting the nuanced challenges of real‑world voice interactions.
By Pritish Mishra, Ishaan Kumar, Akshat Mandoli, Sudarshan Kamath
The paper introduces ContraTalk, a benchmark that tests whether dialogue models truly use acoustic cues or rely on transcript shortcuts. It formalizes cross‑modal disagreement, creates conflict and consistent QA examples, and proposes an Audio Twin representation to expose acoustic evidence to models. Experiments show that while text‑only LLMs perform well on consistent cases, they falter on conflict cases, and AudioLLMs only partially mitigate this issue.
By Yen-Ju Lu, Yuzhe Wang, Yaohan Guan, Xiluo He, Jiarui Hai, Mingrui Liang, Kaavya Chaparala, Thomas Thebaud, Laureano Moro-Velazquez, Najim Dehak, Jesus Villalba
arXiv:2607. 16610v1 Announce Type: new Abstract: Long-horizon AI agents are becoming increasingly capable, yet their interaction with users remains surprisingly thin.
By Chen Chen, Zhehuai Chen
arXiv:2608. 19515v1 Announce Type: new Abstract: Prosodic cues can convey task-relevant information that alters the trajectory and outcome of a task-oriented dialogue, even when the words themselves remain unchanged.
By Xinyi Liu, Hooshang Nayyeri, Dilek Hakkani-Tur, Emine Yilmaz, JK Kim, Yifei Zhang, Charith Peris, Hari Thadakamalla