arXiv:2609.14666v1 Announce Type: cross
Abstract: Turn-taking is a fundamental component of spoken interaction, and while humans naturally rely on both verbal and non-verbal signals, dialogue systems...
By Willem Berner, Julio Cesar Cavalcanti, Kalle {\AA}str\"om, Gabriel Skantze
arXiv:2606. 16731v1 Announce Type: cross Abstract: Current multiparty turn-taking models often rely on complex microphone arrays or multi-camera setups, limiting their applicability in human-robot interaction scenarios.
By Haotian Qi, Gabriel Skantze
arXiv:2606. 16568v1 Announce Type: cross Abstract: Reliable turn-taking is essential for spoken dialogue systems.
By Rutherford A. Patamia, Ming Liu, Wei Luo, Favour Ekong, Akan Cosgun
arXiv:2609.10394v1 Announce Type: cross
Abstract: Current audio-visual speech recognition (AVSR) benchmarks, like LRS3, rely heavily on clean, scripted and rehearsed speech. They fail to reflect the...
By Rishabh Jain, Aristeidis Papadopoulos, Zhaofeng Lin, Naomi Harte
TurnBench is a new multi‑domain benchmark for evaluating turn‑taking dynamics in spoken dialogue. It comprises a 30‑hour hand‑labeled corpus of dyadic human conversations, a standardized evaluation protocol for end‑of‑turn and interruption detection, and covers six distinct interaction styles with triple annotation. The benchmark also provides a 104‑hour training set, a public leaderboard, and an interactive dataset viewer at https://turnbench.sesame.com.
By Freeman Jiang, Ramon Sanabria, Soham Deshmukh, Bandhav Veluri, Simon Michael Vuch Williams, Elliott K. Suen, Garreth Lee, Kevin Yoonho Choi, Takuya Umeki, Riku Kubo, Sathvik Udupa, Chien-yu Huang, Shih-Yun Shan Kuan, Zhuoyan Tao, Satyapriya Krishna, Sefik Emre Eskimez, Yu Tsao, Hung-yi Lee, Shinji Watanabe
Turn-taking prediction is a key requirement for social robots involved in human-human interaction, particularly in mediator settings, where the robot must anticipate conversational dynamics rather than merely react to pauses. This work presents a Multimodal Voice Activity Projection (MM-VAP) framework that extends the original audio-only VAP formulation to synchronized audio-visual inputs while preserving its self-supervised future-projection objective.
The paper introduces AV-STE, a modular streaming audio‑visual front‑end that enhances corrupted semantic speech tokens using noisy audio and lip video before they reach a frozen speech LLM. By preserving the downstream dialogue model’s pretrained conversational abilities, AV‑STE improves response coherence from 1.42 to 1.91 in same‑dataset speaker interference scenarios while maintaining turn‑taking behavior. These gains also transfer to out‑of‑domain Seamless Interaction.
By Bella Godiva, Yeonju Kim, Yong Man Ro
arXiv:2508. 08237v4 Announce Type: replace-cross Abstract: The emergence of audio-visual foundation models underscores the importance of reliably assessing their multi-modal understanding.
By Daniil Zverev, Thadd\"aus Wiedemer, Ameya Prabhu, Matthias Bethge, Wieland Brendel, A. Sophia Koepke
arXiv:2606. 15751v1 Announce Type: cross Abstract: Audio-Language Models (ALMs) have shown remarkable success in zero-shot audio classification by aligning audio waveforms with text.
By Hyebin Cho, Jaehyuk Jang, Changick Kim, Joon Son Chung
arXiv:2607. 07294v1 Announce Type: cross Abstract: Turn-taking prediction is a key requirement for social robots involved in human-human interaction, particularly in mediator settings, where the robot must anticipate conversational dynamics rather than merely react to pauses.
By Antonio Cano, Guillermo P\'erez, Luis Merino, Randy Gomez
Existing multi-speaker dialogue systems bind speakers to utterances through structured supervision: per-turn tags, multi-stream transcriptions, or learnable speaker embeddings. These systems operate within speech-only pipelines that produce clean vocal sequences without the ambient texture of real conversations.
arXiv:2604.14129v2 Announce Type: replace
Abstract: While Audio-Visual Language Models (AVLMs) have achieved remarkable progress over recent years, their reliability is bottlenecked by cross-modal ha...
By Ami Baid, Zihui Xue, Kristen Grauman