arXiv AI By David Nadrchal, Monorama Swain, Florian Schmid, Gerhard Widmer, Paul Primus

Personalized Automatic Speech Recognition for a Dysarthric and Tracheostomic Speaker using Artificial Conversations

Read the original on arXiv AI →

This paper introduces a personalized automatic speech recognition system for a Czech speaker with a permanent tracheal stoma and severe dysarthria. The authors release a 33‑hour annotated dataset collected via an artificial conversation protocol and develop a multi‑stage training pipeline based on Whisper Base, fine‑tuning on Czech speech, simulated tracheostomic speech, and the speaker’s data. Evaluations in scripted, question‑answering, and spontaneous dialogue scenarios show a 50 % relative reduction in character error rate compared to the Whisper Base baseline and better accuracy than the speaker’s assistants on isolated utterances.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jun 2

Efficient ASR Training with Conversations that Never Happened

Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that generates scenario-level dialogues with participant metadata, maps speaker attributes to TTS voice profiles, and assembles synthesized utterances into speaker-aware simulated conversations.

Hugging Face Trending Papers
Jul 20

Re-Sonance: A Dysarthric Asynchronous Real-Time Speech Conversion System Based on a Three-Stage Cascaded ASR-LLM-TTS Architecture

Individuals with dysarthria face significant challenges in professional speaking scenarios such as conferences, presentations, and meetings, where real-time communication is crucial. While existing Augmentative and Alternative Communication (AAC) systems provide basic support, they often fail to meet the demands of professional speaking environments due to high latency and unnatural speech patterns.

arXiv AI
Aug 28

From Sound to Symptom: Real-Time Respiratory Signal Understanding for Conversational Healthcare Agents

The paper introduces HealthCUES, a real‑time streaming pipeline that extracts and analyzes cough and throat‑clearing events from live spoken conversations. It detects coughs within sub‑second latency, distinguishes cough subtypes (dry, wet, barking, whooping), differentiates coughing from throat clearing, and estimates temporal boundaries, all while gating alerts based on conversational context. The system, built on Qwen3Omni, achieves high accuracy (93% F1 for cough detection) and low latency (340 ms) and has been validated by healthcare professionals for telehealth use.

By Tanmay Laud, Herprit Mahal, Subhabrata Mukherjee
arXiv Computation and Language
Sep 22

The Bairong System for MLC-SLM 2026: Dynamic Question-Aware Evidence Routing for Multilingual Conversational Speech Understanding

arXiv:2609.22214v1 Announce Type: new Abstract: Long multilingual conversational spoken question answering requires systems to balance long-range transcript semantics with sparse acoustic and speaker...

By Shangkun Huang, Junchao Hu, Huan Shen, Guoji Wang, Yingao Wang, Shaosai Li, Wei Zou, Yunzhang Chen