VoxReason introduces a listener‑free evaluation framework that measures whether a speech planning system’s delivery choices—such as pitch, energy, rate, pause, emphasis, and stance—are grounded in cited source records before any waveform is generated. The system outputs a source‑cited speaking plan and uses a deterministic verifier to check citation legality, slot agreement, unsupported states, schema validity, and counterfactual locality. Experiments on 1,440 source‑label cases show that simple slot accuracy can be misleading, while a 7B locality‑based repair model significantly improves plan‑slot accuracy and locality, and removing source records sharply reduces the grounded score.
whyItMatters":"The framework provides a concrete, measurable way to ensure that expressive speech systems make source‑licensed planning decisions, addressing a source‑use failure that occurs before audio synthesis."
By Mengzhe Geng
arXiv:2609.38232v1 Announce Type: cross
Abstract: Streaming spoken agents may take an external action before the available speech supports it, yet final-turn scores do not reveal whether each observe...
By Mengzhe Geng
The paper introduces Inquesto Score (IS), a protocol that measures voice‑agent reliability by calculating the percentage of calls that reach the caller’s goal without functional failure. IS defines explicit failure events and severity levels, evaluates timing, semantic, and state‑dependent failures using audio, scenario predicates, tool traces, and a pinned open‑model judge, and provides diagnostic views on behavior, acoustic robustness, identity handling, and speaker groups. The authors evaluate IS v0.1 on 30 scenarios, three acoustic conditions, four speaker groups, and 306 calls per agent across 13 configurations, demonstrating that reliable measurement requires evidence beyond transcripts and explicit treatment of deployment conditions.
By Massa Baali, Bhiksha Raj
arXiv:2608.29241v1 Announce Type: new
Abstract: Clinical voice agents are now deployed in routine care, where real patients do not wait their turn: they interrupt. These systems typically use a casca...
By Zachary Ellis, Spencer Hazel, Adam Brandt, Yajie Vera He, Ernest Lim, Jared Joselowitz
arXiv:2608. 19515v1 Announce Type: new Abstract: Prosodic cues can convey task-relevant information that alters the trajectory and outcome of a task-oriented dialogue, even when the words themselves remain unchanged.
By Xinyi Liu, Hooshang Nayyeri, Dilek Hakkani-Tur, Emine Yilmaz, JK Kim, Yifei Zhang, Charith Peris, Hari Thadakamalla
arXiv:2608.29136v1 Announce Type: cross
Abstract: We ask whether a model protects a user in the same way when that user speaks rather than types. Using a single distress vignette---a physical injury...
By Eunna Lee, Soomyoung Lee, Jungpyo Nam, Heonjin Ha, Jamin Jung, Kyunam Choi, Sunjun Hwang, Yeonghun Kim, Seok-Jae Lim
Full‑duplex speech models can listen and speak simultaneously, but they struggle to decide when to speak. Experiments with five model families show that being addressed or encountering silence are reliable triggers, whereas cues like false facts or hazards are not. Even when models answer questions, they rarely challenge false claims or warn about danger, revealing a gap in content understanding and intervention decisions.
By Linkai Peng, Baorian Nuchged, Kaiqi Fu, Yuyang Yao
arXiv:2609.03203v3 Announce Type: replace-cross
Abstract: Speech systems increasingly infer how an utterance should be delivered from context, but a plausible delivery plan may not be supported by th...
By Mengzhe Geng
arXiv:2609.38512v1 Announce Type: new
Abstract: Voice agents in production must handle several requests, background speech, and customers who lose patience. We introduce VAmoS Energy, a benchmark tha...
By Joshua Meyer, Sahar Shayegan, Ritiz Tambi, Ali Khan, Sun Kim, Victor Shih, Mehdi Jamei, Andi Partovi
arXiv:2609.30483v1 Announce Type: cross
Abstract: Audio language models state numbers for acoustic quantities, and neither human opinion nor a judge model says whether such a number is true of the si...
By Sheng-Tse Lin, Siyuan Zhai, Chien-Liang Kuo, Massa Baali, Bhiksha Raj
Prosodic cues can convey task-relevant information that alters the trajectory and outcome of a task-oriented dialogue, even when the words themselves remain unchanged. Yet existing benchmarks typically evaluate prosodic perception, response appropriateness, and task-oriented dialogue in isolation, making it difficult to test whether prosodic evidence changes downstream decisions.
arXiv:2608.28916v1 Announce Type: new
Abstract: Automatic speech recognition (ASR) systems are commonly evaluated with word error rate (WER), yet many voice workflows depend on exact written values f...
By Tyler Baumgartner, Brandon Tai, Lisa Kaelin-Martin, Candice Fan, Luc Debaupte, Bill Wang, Yi Zhong