arXiv AI

Personalized Automatic Speech Recognition for a Dysarthric and Tracheostomic Speaker using Artificial Conversations

This paper introduces a personalized automatic speech recognition system for a Czech speaker with a permanent tracheal stoma and severe dysarthria. The authors release a 33‑hour annotated dataset collected via an artificial conversation protocol and develop a multi‑stage training pipeline based on Whisper Base, fine‑tuning on Czech speech, simulated tracheostomic speech, and the speaker’s data. Evaluations in scripted, question‑answering, and spontaneous dialogue scenarios show a 50 % relative reduction in character error rate compared to the Whisper Base baseline and better accuracy than the speaker’s assistants on isolated utterances.

Hugging Face Trending Papers
Jun 2

Efficient ASR Training with Conversations that Never Happened

Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that generates scenario-level dialogues with participant metadata, maps speaker attributes to TTS voice profiles, and assembles synthesized utterances into speaker-aware simulated conversations.

Hugging Face Trending Papers
Jul 20

Re-Sonance: A Dysarthric Asynchronous Real-Time Speech Conversion System Based on a Three-Stage Cascaded ASR-LLM-TTS Architecture

Individuals with dysarthria face significant challenges in professional speaking scenarios such as conferences, presentations, and meetings, where real-time communication is crucial. While existing Augmentative and Alternative Communication (AAC) systems provide basic support, they often fail to meet the demands of professional speaking environments due to high latency and unnatural speech patterns.

arXiv AI
Aug 28

From Sound to Symptom: Real-Time Respiratory Signal Understanding for Conversational Healthcare Agents

The paper introduces HealthCUES, a real‑time streaming pipeline that extracts and analyzes cough and throat‑clearing events from live spoken conversations. It detects coughs within sub‑second latency, distinguishes cough subtypes (dry, wet, barking, whooping), differentiates coughing from throat clearing, and estimates temporal boundaries, all while gating alerts based on conversational context. The system, built on Qwen3Omni, achieves high accuracy (93% F1 for cough detection) and low latency (340 ms) and has been validated by healthcare professionals for telehealth use.

By Tanmay Laud, Herprit Mahal, Subhabrata Mukherjee
arXiv Computation and Language
Sep 22

The Bairong System for MLC-SLM 2026: Dynamic Question-Aware Evidence Routing for Multilingual Conversational Speech Understanding

arXiv:2609.22214v1 Announce Type: new Abstract: Long multilingual conversational spoken question answering requires systems to balance long-range transcript semantics with sparse acoustic and speaker...

By Shangkun Huang, Junchao Hu, Huan Shen, Guoji Wang, Yingao Wang, Shaosai Li, Wei Zou, Yunzhang Chen
arXiv Computation and Language
2d ago

An automated pipeline for standardised speech-unit annotation in spontaneous dialogue

The paper introduces an automated pipeline that extracts conversational turns and backchannels from separate-channel recordings of spontaneous dyadic dialogue, combining voice activity detection, channel-energy filtering, temporal merging, automatic speech recognition, and context-based post‑processing. Evaluated on 99 ten‑minute Danish conversations, the system achieved F1 scores around 0.62 for both turns and backchannels, with median onset/offset errors of roughly 0.15–0.18 s. Performance was consistent across normal and asymmetric listening conditions, and a case study showed the pipeline’s outputs were less variable than human annotations, supporting its use as a reliable first‑pass annotation tool in semi‑automated workflows.

By Hanlu He, Harald Vilhelm Skat-R{\o}rdam, Ingvi \"Orn\'olfsson, Ivana Konvalinka
arXiv Computation and Language
Sep 28

Inference-Time Target Speaker Unlearning in LLM-Based Automatic Speech Recognition

The paper introduces a new target‑speaker unlearning task for automatic speech recognition (TSU‑ASR) that allows certain speakers to opt out of transcription while still indicating their presence. A lightweight Enrollment‑Conditioned Gating (ECG) module is added to a frozen dual‑stream speech LLM, enabling dynamic unlearning of new opt‑out speakers during inference. Experiments on AMI and AliMeeting datasets show significant drops in transcription accuracy for opt‑out speakers while preserving performance for retained speakers.

By Bo Su, Yueru Yan, Thai Le
arXiv AI
Jul 17

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems

arXiv:2607. 14846v1 Announce Type: cross Abstract: Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual representation.

By David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer, Jakub Piotr C{\l}apa, Theo Lebryk, Jens Madsen, Olya Ossipova, Sharath Rao, Hoon Shin, Tigran Soghbatyan, Georg Streich, Rashish Tandon, Panagiotis Tzirakis
arXiv Machine Learning
Jun 8

SEAM: Shortcut-Aware Real-Time Detection of Scripted vs. Spontaneous Speech for Interview Guardrails

arXiv:2606. 06837v1 Announce Type: cross Abstract: Scripted vs spontaneous speech detection is appealing for interview guardrails, but benchmark performance can be inflated by shortcuts tied to corpus identity, channel conditions, and recording artifacts rather than speaking style itself.

By Vsevolod (V.), Kovalev, Pranay Manocha