arXiv AI By Hao Zhang, Thomas Thebaud, Georgi Tinchev, Venkatesh Ravichandran, Laureano Moro-Velazquez

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue

Read the original on arXiv AI →

arXiv:2607. 01345v1 Announce Type: cross Abstract: Turn-taking naturalness is central to full-duplex spoken dialogue systems, yet its automatic evaluation remains limited.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 27

TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue

TurnBench is a new multi‑domain benchmark for evaluating turn‑taking dynamics in spoken dialogue. It comprises a 30‑hour hand‑labeled corpus of dyadic human conversations, a standardized evaluation protocol for end‑of‑turn and interruption detection, and covers six distinct interaction styles with triple annotation. The benchmark also provides a 104‑hour training set, a public leaderboard, and an interactive dataset viewer at https://turnbench.sesame.com.

By Freeman Jiang, Ramon Sanabria, Soham Deshmukh, Bandhav Veluri, Simon Michael Vuch Williams, Elliott K. Suen, Garreth Lee, Kevin Yoonho Choi, Takuya Umeki, Riku Kubo, Sathvik Udupa, Chien-yu Huang, Shih-Yun Shan Kuan, Zhuoyan Tao, Satyapriya Krishna, Sefik Emre Eskimez, Yu Tsao, Hung-yi Lee, Shinji Watanabe
arXiv Computation and Language
Sep 25

Same Words, Different Actions: Paired Turn-Taking Evaluation under Rewritten Dialogue Contexts

The paper introduces ECHO, a paired diagnostic benchmark for evaluating Chinese real‑time spoken dialogue systems on turn‑taking. ECHO pairs examples that share the same overlap transcript but differ in preceding multi‑turn context, requiring either Yield or Keep actions, and also includes off‑talk cases to test unnecessary yielding. Experiments on four speech systems reveal that three systems over‑yield, correctly keeping the floor on fewer than 13% of backchannels, while the fourth system shows a more balanced performance, illustrating that interruption‑only evaluation can overestimate turn‑taking reliability.

By Shuofeng Zhao, Hongwei Cai, Wenke Fan, Qingxiang Guo, Zhou Wang, Dawei Yang, Zhiyang Zhou, Yingxin Shang, Weixu Wang, Lin Yang, Shuran Zhou, Yang Song
arXiv Computation and Language
Sep 11

Using Semantic Uncertainty to Estimate Transition Relevance in Turn-taking

The paper proposes using semantic uncertainty, derived from large language models, to predict Transition Relevance Places (TRPs) in spoken dialogue. By sampling possible continuations of an ongoing turn and measuring changes in semantic dispersion, the authors identify moments when a listener might take the floor. Their method outperforms prompt-based and fine-tuned text-only baselines on a dataset with real-time TRP labels, supporting the idea that evolving semantic constraints inform turn‑taking opportunities in unscripted interaction.

By Muhammad Umair, Jan P. de Ruiter
arXiv Computation and Language
Sep 1

DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues

arXiv:2607.26178v2 Announce Type: replace Abstract: Turn-taking is a central component of full-duplex interaction. Which turn-taking behaviors are appropriate varies with the scenario, yet current mo...

By Takyoung Kim, Kang-wook Kim, Sang Hoon Woo, Julia Hirschberg, Gunhee Kim, Dilek Hakkani-T\"ur