arXiv AI

Fast When, Careful Who: Dual-Process Multiparty Turn-Taking with Diffusion Augmentation

arXiv:2606. 16568v1 Announce Type: cross Abstract: Reliable turn-taking is essential for spoken dialogue systems.

arXiv Computation and Language
Sep 16

Audio-Visual Turn-taking Prediction in Cocktail Party Scenarios

The paper evaluates audio‑visual predictive turn‑taking models trained on clean data when applied to a noisy cocktail‑party scenario derived from the AVCocktail dataset. Results show a consistent performance drop—up to 38% relative in weighted F1—across both audio and visual modalities, with fine‑tuning improving robustness but varying by modality and pre‑training data size. The study highlights differing generalisation and adaptation abilities of audio versus visual inputs and underscores the need for robust modelling strategies in noisy human interactions.

By Long-Vu Hoang, Naomi Harte
arXiv Computation and Language
Aug 27

TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue

TurnBench is a new multi‑domain benchmark for evaluating turn‑taking dynamics in spoken dialogue. It comprises a 30‑hour hand‑labeled corpus of dyadic human conversations, a standardized evaluation protocol for end‑of‑turn and interruption detection, and covers six distinct interaction styles with triple annotation. The benchmark also provides a 104‑hour training set, a public leaderboard, and an interactive dataset viewer at https://turnbench.sesame.com.

By Freeman Jiang, Ramon Sanabria, Soham Deshmukh, Bandhav Veluri, Simon Michael Vuch Williams, Elliott K. Suen, Garreth Lee, Kevin Yoonho Choi, Takuya Umeki, Riku Kubo, Sathvik Udupa, Chien-yu Huang, Shih-Yun Shan Kuan, Zhuoyan Tao, Satyapriya Krishna, Sefik Emre Eskimez, Yu Tsao, Hung-yi Lee, Shinji Watanabe
Hugging Face Trending Papers
Jul 8

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders

Turn-taking prediction is a key requirement for social robots involved in human-human interaction, particularly in mediator settings, where the robot must anticipate conversational dynamics rather than merely react to pauses. This work presents a Multimodal Voice Activity Projection (MM-VAP) framework that extends the original audio-only VAP formulation to synchronized audio-visual inputs while preserving its self-supervised future-projection objective.

arXiv AI
Jul 9

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders

arXiv:2607. 07294v1 Announce Type: cross Abstract: Turn-taking prediction is a key requirement for social robots involved in human-human interaction, particularly in mediator settings, where the robot must anticipate conversational dynamics rather than merely react to pauses.

By Antonio Cano, Guillermo P\'erez, Luis Merino, Randy Gomez
arXiv AI
Sep 12

Less can be More: What Aspects of Speech Drive End-of-Turn Detection

The study investigates which speech aspects best signal the end of a speaker’s turn in conversational AI. By systematically ablating acoustic, prosodic, and semantic cues in a lightweight trimodal classifier, the authors find that combining acoustic and prosodic features yields the highest accuracy and lowest latency, achieving an utterance F1 of 0.93 with 7.8% false alarms at 400 ms median latency. Adding textual information actually increases premature detections without improving performance, and feature analysis shows prosodic cues provide the strongest class separability while text representations overlap significantly.

By Rini Sharon, Manickavela A, Kadri Hacioglu, Andreas Stolcke
arXiv AI
Jun 18

Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors

arXiv:2606. 19325v1 Announce Type: cross Abstract: Existing multi-speaker dialogue systems bind speakers to utterances through structured supervision: per-turn tags, multi-stream transcriptions, or learnable speaker embeddings.

By Michael Finkelson, Daniel Segal, Eitan Richardson, Shahar Armon, Nani Goldring, Poriya Panet, Nir Zabari, Benjamin Brazowski, Or Patashnik, Yoav HaCohen
Hugging Face Trending Papers
Jun 11

Adaptive Turn-Taking for Real-time Multi-Party Voice Agents

Turn-taking in multi-party spoken conversations remains a fundamental challenge for voice-based agents, particularly under dynamic floor competition and varying user expectations. We propose ModeratorLM, a role-playing voice agent that conditions turn-taking behavior on an explicitly assigned role in multi-party settings.