Hy‑MultiTurn is a Chinese benchmark designed to evaluate deep multi‑turn dialogue understanding over long interactions. It introduces six controlled evaluation modes—constraint memory, precise execution, constraint synthesis, object localization, action suppression, and reference resolution—across 209 tasks ranging from 12 to 76 turns, incorporating dialogue length, irrelevant distractions, and colloquial phrasing. Testing 22 state‑of‑the‑art models shows the benchmark is highly challenging, with even the best model meeting all criteria only 41.1% of the time and no model excelling in every mode.
By Eileen Ye, Jiawen Tao, Yaoming Li, Chenxu Liu, Wenhan Yu, Yaxin Fan, Xiaokun Yuan, Mengzhou Wu, Yanbing Jiang, Maxm Pan
arXiv:2609.35812v1 Announce Type: new
Abstract: In-car conversational assistants (ICAs) are increasingly integrated into vehicles to support route planning, vehicle control, and information access. E...
By Vaishnav Negi, Lev Sorokin, Soroosh Tayebi Arasteh, Andrea Stocco
The paper introduces ECHO, a paired diagnostic benchmark for evaluating Chinese real‑time spoken dialogue systems on turn‑taking. ECHO pairs examples that share the same overlap transcript but differ in preceding multi‑turn context, requiring either Yield or Keep actions, and also includes off‑talk cases to test unnecessary yielding. Experiments on four speech systems reveal that three systems over‑yield, correctly keeping the floor on fewer than 13% of backchannels, while the fourth system shows a more balanced performance, illustrating that interruption‑only evaluation can overestimate turn‑taking reliability.
By Shuofeng Zhao, Hongwei Cai, Wenke Fan, Qingxiang Guo, Zhou Wang, Dawei Yang, Zhiyang Zhou, Yingxin Shang, Weixu Wang, Lin Yang, Shuran Zhou, Yang Song
CallScreenBench is a benchmark for evaluating small, on-device language models that act as phone secretaries, focusing on their ability to handle unknown inbound calls without a cooperative task. The benchmark measures owner endorsement through five call-and-note metrics, each paired with counter-metrics and uncertainty estimates, and includes guardedness diagnostics to identify safe, tool‑free proxies. Results across 4‑bit checkpoints of 0.6‑4 B parameter models show varying performance on service, recall, plausibility, and triage discrimination, highlighting trade‑offs between quality and guardedness.
By Jiaqi Gan, Haoyuan Tang, Jamey Z. Liang, Siying Chen, Ankit Raj, Kidus Zewde, Yuchen Zhou, Yuxin Zhang, Simiao Ren
arXiv:2607.26178v2 Announce Type: replace
Abstract: Turn-taking is a central component of full-duplex interaction. Which turn-taking behaviors are appropriate varies with the scenario, yet current mo...
By Takyoung Kim, Kang-wook Kim, Sang Hoon Woo, Julia Hirschberg, Gunhee Kim, Dilek Hakkani-T\"ur
arXiv:2604. 03924v2 Announce Type: replace-cross Abstract: Goal-oriented conversational systems require making sequential decisions under uncertainty about the user's intent, where the algorithm must balance information acquisition and target commitment over multiple turns.
By Xinyi Ling, Ye Liu, Reza Averly, Xia Ning
arXiv:2606. 13544v1 Announce Type: cross Abstract: Turn-taking in multi-party spoken conversations remains a fundamental challenge for voice-based agents, particularly under dynamic floor competition and varying user expectations.
By Soumyajit Mitra, Prabhat Pandey, Abhinav Jain, Shanmukha Sahith, K V Vijay Girish
arXiv:2609.22256v1 Announce Type: new
Abstract: Effective workplace negotiation requires balancing multiple objectives, including achieving task goals, preserving professional relationships, and reso...
By Bibhuti Jha, Rishikant Chigrupaatii, Priyanshu Priya, Asif Ekbal
The paper introduces a pipeline that generates intent‑labeled, two‑channel conversational speech from relational event lists, enabling controlled synthesis of full‑duplex dialogue with 42 phenomena across eight families in English and Mandarin. By having an LLM author each event’s speaker, text, conversational act, and attachment, and then aligning and timing these events independently, the system produces diverse, realistic turn‑taking signals. Experiments show that models trained on this synthetic corpus achieve higher floor‑occupancy accuracy and better start‑speaking/listening F1 scores compared to models trained on prior data.
By Matthew Sun, Vinay Kothapally, Meng Yu, Chao Huang, Hao Zhang, Yixuan Zhang, Steve Yves
TurnBench is a new multi‑domain benchmark for evaluating turn‑taking dynamics in spoken dialogue. It comprises a 30‑hour hand‑labeled corpus of dyadic human conversations, a standardized evaluation protocol for end‑of‑turn and interruption detection, and covers six distinct interaction styles with triple annotation. The benchmark also provides a 104‑hour training set, a public leaderboard, and an interactive dataset viewer at https://turnbench.sesame.com.
By Freeman Jiang, Ramon Sanabria, Soham Deshmukh, Bandhav Veluri, Simon Michael Vuch Williams, Elliott K. Suen, Garreth Lee, Kevin Yoonho Choi, Takuya Umeki, Riku Kubo, Sathvik Udupa, Chien-yu Huang, Shih-Yun Shan Kuan, Zhuoyan Tao, Satyapriya Krishna, Sefik Emre Eskimez, Yu Tsao, Hung-yi Lee, Shinji Watanabe
arXiv:2603.16783v2 Announce Type: replace
Abstract: Robust voice agents require exposure to the full diversity of how people interact through speech. However, obtaining enough spoken interactions is...
By Jonggeun Lee, Junseong Pyo, Jeongmin Park, Yohan Jo
arXiv:2608. 15755v1 Announce Type: new Abstract: User-centric multi-turn agents must act on an evolving task situation shaped by changing user intents, accumulated tool-grounded facts, missing information, and execution constraints.
By Meiling Tao, Yiling Tao, Peng Wang