arXiv:2607. 07294v1 Announce Type: cross Abstract: Turn-taking prediction is a key requirement for social robots involved in human-human interaction, particularly in mediator settings, where the robot must anticipate conversational dynamics rather than merely react to pauses.
By Antonio Cano, Guillermo P\'erez, Luis Merino, Randy Gomez
arXiv:2606. 16731v1 Announce Type: cross Abstract: Current multiparty turn-taking models often rely on complex microphone arrays or multi-camera setups, limiting their applicability in human-robot interaction scenarios.
By Haotian Qi, Gabriel Skantze
arXiv:2609.14666v1 Announce Type: cross
Abstract: Turn-taking is a fundamental component of spoken interaction, and while humans naturally rely on both verbal and non-verbal signals, dialogue systems...
By Willem Berner, Julio Cesar Cavalcanti, Kalle {\AA}str\"om, Gabriel Skantze
arXiv:2609.13076v1 Announce Type: cross
Abstract: Conversational voice agents have advanced significantly, offering increasingly natural human-machine interactions through both cascaded and end-to-en...
By Yi-Jen Shih, Shih-Yun Shan Kuan, Guan-Ting Lin, Kai-Wei Chang, Siddhant Arora, Shu-wen Yang, Abdelrahman Mohamed, Shinji Watanabe, Hung-yi Lee, David Harwath
arXiv:2606. 16568v1 Announce Type: cross Abstract: Reliable turn-taking is essential for spoken dialogue systems.
By Rutherford A. Patamia, Ming Liu, Wei Luo, Favour Ekong, Akan Cosgun
Turn-taking in multi-party spoken conversations remains a fundamental challenge for voice-based agents, particularly under dynamic floor competition and varying user expectations. We propose ModeratorLM, a role-playing voice agent that conditions turn-taking behavior on an explicitly assigned role in multi-party settings.
arXiv:2607.26178v2 Announce Type: replace
Abstract: Turn-taking is a central component of full-duplex interaction. Which turn-taking behaviors are appropriate varies with the scenario, yet current mo...
By Takyoung Kim, Kang-wook Kim, Sang Hoon Woo, Julia Hirschberg, Gunhee Kim, Dilek Hakkani-T\"ur
The article surveys multi‑turn conversational AI, highlighting its shift from isolated text prompts to sustained, multimodal interactions that involve clarifying goals, revising requests, and switching topics. It reviews literature across text‑only dialogue, AudioLLMs, multimodal and omni‑modal systems, and tool‑augmented agents, organizing findings around datasets, models, training, evaluation, and cross‑cutting challenges. The analysis reveals that while multimodal perception and action have progressed rapidly, systems still struggle with persistent memory, cross‑turn grounding, full‑duplex interaction, robust evaluation, and cultural alignment.
By Syeda Faiza Ahmed, Zien Sheikh Ali, Hunzalah Hassan Bhatti, Firoj Alam, Shammur Absar Chowdhury
arXiv:2606. 13544v1 Announce Type: cross Abstract: Turn-taking in multi-party spoken conversations remains a fundamental challenge for voice-based agents, particularly under dynamic floor competition and varying user expectations.
By Soumyajit Mitra, Prabhat Pandey, Abhinav Jain, Shanmukha Sahith, K V Vijay Girish
ECHO-G is a framework for generating full‑body co‑speech motion for humanoid robots, jointly conditioned on speech audio and timed transcripts. Its Speech‑Grounded Diffusion Transformer (SGDiT) fuses frame‑aligned acoustic features with token‑level linguistic context, preserving distinct granularities while modeling one‑to‑many utterance‑motion relationships directly in robot space. The authors introduce a BEAT2‑derived audio‑text‑robot dataset, a benchmark for co‑speech characteristics, robot‑motion quality, and runtime efficiency, and demonstrate that direct robot‑space generation outperforms human‑motion generation and retargeting pipelines, with joint audio‑text conditioning yielding superior results in both quantitative evaluation and a video‑rating study.
"whyItMatters":"The study provides a new dataset, benchmark, and a demonstrably effective method for generating realistic co‑speech motion directly in robot space, advancing practical humanoid robot interaction."
By Yizhao Li, Pusen Gao, Ming Wang, Shaojie Shen, Shuo Yang, Hao Xu
Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework jointly...
Existing multi-speaker dialogue systems bind speakers to utterances through structured supervision: per-turn tags, multi-stream transcriptions, or learnable speaker embeddings. These systems operate within speech-only pipelines that produce clean vocal sequences without the ambient texture of real conversations.