arXiv AI By Yuqi Wang, Fengyuan Liu, Haochen Luo, Zhiqi Yu, Qi Liu

RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken Dialogue

Read the original on arXiv AI →

RoleBreak is an open benchmark designed to evaluate long‑horizon role‑playing robustness in spoken dialogue systems. It includes 310 character‑based and user‑centered roles, 6,688 human‑verified dialogue turns, and 11,743 fine‑grained evaluation criteria, with 1,856 turns specifically targeting expressive vocal emotion. The benchmark stresses role consistency, interaction quality, safety, and affect over extended conversations, and the authors evaluated nine system configurations across full‑duplex, omni‑modal, and cascaded ASR–LLM–TTS paradigms.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 17

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems

arXiv:2607. 14846v1 Announce Type: cross Abstract: Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual representation.

By David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer, Jakub Piotr C{\l}apa, Theo Lebryk, Jens Madsen, Olya Ossipova, Sharath Rao, Hoon Shin, Tigran Soghbatyan, Georg Streich, Rashish Tandon, Panagiotis Tzirakis