arXiv:2609.38867v1 Announce Type: new
Abstract: Large language model (LLM) computer-use agents are typically evaluated with clean written instructions, despite speech being an increasingly popular in...
By Terumi Chiba, Guangzhi Sun, Zheqi Yuan, Chao Zhang
The paper evaluates the use of large language models (LLMs) as judges for assessing conversational voice agents, comparing human judgments with GPT‑4.1 and GPT‑5 across telecom and retail interactions. It examines agreement, metric‑level correlations, and consistency across three evaluation configurations (p0, p1, p2) to determine how reliably LLMs can judge conversational quality and safety. The study finds that LLM‑based evaluation can be effective but its reliability varies by metric and configuration, suggesting a hybrid approach where LLMs handle scalable assessment while humans focus on metrics requiring contextual interpretation.
By Anupam Purwar, Shashank Singh, Kritika Srivastava
arXiv:2609.24812v1 Announce Type: new
Abstract: Voice provides a natural and immediate interface for AI agents. Many settings in which voice agents could be useful, including meetings, households, an...
By Chenxu Xiong, Dongming Shen, Yuzhi Tang, Wentao Ma, Mu Li, Alex Smola
arXiv:2603.16783v2 Announce Type: replace
Abstract: Robust voice agents require exposure to the full diversity of how people interact through speech. However, obtaining enough spoken interactions is...
By Jonggeun Lee, Junseong Pyo, Jeongmin Park, Yohan Jo
The paper reviews the fragmented literature on real‑time voice agents, noting that architecture, turn‑taking, and agentic evaluation communities rarely cite each other. It presents three evidence‑based claims: (1) architecture choice is a deployment constraint rather than a definitive solution, (2) evaluation has shifted from component quality to grounded outcomes, and (3) the dyadic assumption is breaking down as multiparty turn‑taking and reasoning become essential. The authors propose the TRG reporting standard to characterize agents by timing, recovery, and state‑verified outcomes, with an optional fourth axis for multiparty contexts.
By Shivam Negi, Arpit Rawat, Rashi Jain
arXiv:2609.13602v1 Announce Type: new
Abstract: Voice agents often need to collect names, addresses, identifiers, dates, and times exactly, yet end-to-end benchmarks obscure where capture fails. We i...
By Soham Ray, Victor Barres
The paper introduces Inquesto Score (IS), a protocol that measures voice‑agent reliability by calculating the percentage of calls that reach the caller’s goal without functional failure. IS defines explicit failure events and severity levels, evaluates timing, semantic, and state‑dependent failures using audio, scenario predicates, tool traces, and a pinned open‑model judge, and provides diagnostic views on behavior, acoustic robustness, identity handling, and speaker groups. The authors evaluate IS v0.1 on 30 scenarios, three acoustic conditions, four speaker groups, and 306 calls per agent across 13 configurations, demonstrating that reliable measurement requires evidence beyond transcripts and explicit treatment of deployment conditions.
By Massa Baali, Bhiksha Raj
arXiv:2609.13076v1 Announce Type: cross
Abstract: Conversational voice agents have advanced significantly, offering increasingly natural human-machine interactions through both cascaded and end-to-en...
By Yi-Jen Shih, Shih-Yun Shan Kuan, Guan-Ting Lin, Kai-Wei Chang, Siddhant Arora, Shu-wen Yang, Abdelrahman Mohamed, Shinji Watanabe, Hung-yi Lee, David Harwath
arXiv:2605.13841v3 Announce Type: replace-cross
Abstract: Voice agents are increasingly deployed across enterprise applications. However, no existing benchmark jointly addresses realistic conversatio...
By Tara Bogavelli, Gabrielle Gauthier Melan\c{c}on, Katrina Stankiewicz, Oluwanifemi Bamgbose, Fanny Riols, Hoang H. Nguyen, Raghav Mehndiratta, Lindsay Devon Brin, Joseph Marinier, Hari Subramani, Anil Madamala, Sridhar Krishna Nemala, Srinivas Sunkara
arXiv:2607. 12085v1 Announce Type: new Abstract: Evaluating retail conversational agents requires methods beyond lexical-overlap metrics to assess intent alignment, factuality, helpfulness, clarity, tone, and overall response quality.
By Niranjan Kumar M, Balaji Nagarajan, Karthik Nair, Faysal Satter, Nithin Surendran
arXiv:2607. 14846v1 Announce Type: cross Abstract: Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual representation.
By David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer, Jakub Piotr C{\l}apa, Theo Lebryk, Jens Madsen, Olya Ossipova, Sharath Rao, Hoon Shin, Tigran Soghbatyan, Georg Streich, Rashish Tandon, Panagiotis Tzirakis