arXiv Computation and Language

An automated pipeline for standardised speech-unit annotation in spontaneous dialogue

The paper introduces an automated pipeline that extracts conversational turns and backchannels from separate-channel recordings of spontaneous dyadic dialogue, combining voice activity detection, channel-energy filtering, temporal merging, automatic speech recognition, and context-based post‑processing. Evaluated on 99 ten‑minute Danish conversations, the system achieved F1 scores around 0.62 for both turns and backchannels, with median onset/offset errors of roughly 0.15–0.18 s. Performance was consistent across normal and asymmetric listening conditions, and a case study showed the pipeline’s outputs were less variable than human annotations, supporting its use as a reliable first‑pass annotation tool in semi‑automated workflows.

arXiv AI
Sep 25

A Harness for Synthesizing Diverse Naturalistic Full-Duplex Conversations

The paper introduces a pipeline that generates intent‑labeled, two‑channel conversational speech from relational event lists, enabling controlled synthesis of full‑duplex dialogue with 42 phenomena across eight families in English and Mandarin. By having an LLM author each event’s speaker, text, conversational act, and attachment, and then aligning and timing these events independently, the system produces diverse, realistic turn‑taking signals. Experiments show that models trained on this synthetic corpus achieve higher floor‑occupancy accuracy and better start‑speaking/listening F1 scores compared to models trained on prior data.

By Matthew Sun, Vinay Kothapally, Meng Yu, Chao Huang, Hao Zhang, Yixuan Zhang, Steve Yves
arXiv AI
Aug 28

From Sound to Symptom: Real-Time Respiratory Signal Understanding for Conversational Healthcare Agents

The paper introduces HealthCUES, a real‑time streaming pipeline that extracts and analyzes cough and throat‑clearing events from live spoken conversations. It detects coughs within sub‑second latency, distinguishes cough subtypes (dry, wet, barking, whooping), differentiates coughing from throat clearing, and estimates temporal boundaries, all while gating alerts based on conversational context. The system, built on Qwen3Omni, achieves high accuracy (93% F1 for cough detection) and low latency (340 ms) and has been validated by healthcare professionals for telehealth use.

By Tanmay Laud, Herprit Mahal, Subhabrata Mukherjee
arXiv Computation and Language
Aug 27

TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue

TurnBench is a new multi‑domain benchmark for evaluating turn‑taking dynamics in spoken dialogue. It comprises a 30‑hour hand‑labeled corpus of dyadic human conversations, a standardized evaluation protocol for end‑of‑turn and interruption detection, and covers six distinct interaction styles with triple annotation. The benchmark also provides a 104‑hour training set, a public leaderboard, and an interactive dataset viewer at https://turnbench.sesame.com.

By Freeman Jiang, Ramon Sanabria, Soham Deshmukh, Bandhav Veluri, Simon Michael Vuch Williams, Elliott K. Suen, Garreth Lee, Kevin Yoonho Choi, Takuya Umeki, Riku Kubo, Sathvik Udupa, Chien-yu Huang, Shih-Yun Shan Kuan, Zhuoyan Tao, Satyapriya Krishna, Sefik Emre Eskimez, Yu Tsao, Hung-yi Lee, Shinji Watanabe
arXiv AI
2d ago

Personalized Automatic Speech Recognition for a Dysarthric and Tracheostomic Speaker using Artificial Conversations

This paper introduces a personalized automatic speech recognition system for a Czech speaker with a permanent tracheal stoma and severe dysarthria. The authors release a 33‑hour annotated dataset collected via an artificial conversation protocol and develop a multi‑stage training pipeline based on Whisper Base, fine‑tuning on Czech speech, simulated tracheostomic speech, and the speaker’s data. Evaluations in scripted, question‑answering, and spontaneous dialogue scenarios show a 50 % relative reduction in character error rate compared to the Whisper Base baseline and better accuracy than the speaker’s assistants on isolated utterances.

By David Nadrchal, Monorama Swain, Florian Schmid, Gerhard Widmer, Paul Primus
arXiv Machine Learning
Sep 22

AVTR-1: Open Stack for Real-Time Interactive Avatars

arXiv:2609.22913v1 Announce Type: cross Abstract: Talking-head and dyadic models now achieve real-time inference, yet fast motion generation alone does not produce an interactive conversation. A live...

By Artem Kravtsov, Dmitrii Ziganshin, Vsevolod Poletaev, Gleb Balitskiy, Anastasia Tikhonova, Egor Burkov, Vadim Lebedev