Scaling Human and G2P Supervision for Robust Phonetic Transcription
arXiv:2606. 16019v1 Announce Type: cross Abstract: Expert phonetic annotation is costly, especially for non-standard dialects and atypical speech.
arXiv:2606. 15540v1 Announce Type: cross Abstract: Pathological speech from patients with neurodegenerative and neuromotor disorders is often acoustically distorted and linguistically fragmented, making pathological speech reconstruction necessary to recover intended textual content from distorted and incomplete speech recordings.
arXiv:2606. 16019v1 Announce Type: cross Abstract: Expert phonetic annotation is costly, especially for non-standard dialects and atypical speech.
The paper evaluates whether audio‑language models can use multimodal clinical context to improve dysarthric speech recognition. Using a benchmark built on the Speech Accessibility Project dataset, the authors test diagnosis labels, clinician ratings, and detailed clinical descriptions as prompts for nine models. They find that these prompts yield negligible or negative effects on word error rate, though fine‑tuning with LoRA and mixed prompt formats reduces WER by 52% and benefits certain subgroups such as Down syndrome and mild‑severity speakers.
arXiv:2608.31035v1 Announce Type: new Abstract: Codec-based text-to-speech (TTS) models make language-model post-training applicable to speech generation, but it remains unclear when learned perceptu...
arXiv:2608.28970v1 Announce Type: cross Abstract: Current TTS systems typically rely on open-loop, single-pass generation and can produce sporadic local prosodic defects, such as misplaced stress, un...
arXiv:2607. 10191v1 Announce Type: cross Abstract: Generative streaming models for Target Speaker Extraction (TSE) commonly exhibit a quality--intelligibility trade-off, wherein naive optimization for perceptual audio quality tends to degrade speech intelligibility, and conversely.
arXiv:2608. 00722v1 Announce Type: cross Abstract: Language model-based text-to-speech (LM-based TTS) remains vulnerable to speech hallucinations that deviate from the target text.
arXiv:2609.01246v1 Announce Type: new Abstract: Current Large Language Models (LLMs) are primarily optimized for written text, often producing outputs that are grammatically correct and helpful yet p...
arXiv:2511. 13300v1 Announce Type: cross Abstract: Generative models have shown remarkable performance in speech enhancement (SE), achieving superior perceptual quality over traditional discriminative approaches.
arXiv:2607. 23977v1 Announce Type: cross Abstract: Acoustic biomarkers show promise for detecting Alzheimer's Disease (AD), yet whether the cues driving diagnostic AI align with those salient to human listeners is underexplored across languages and genders, where pathological markers and perceptual strategies differ.
The paper investigates how Spoken Language Models (SLMs) process speech compared to text, noting that current SLMs show weak alignment between speech and text representations despite strong downstream performance. The authors propose a framework that separates length mismatch from semantic alignment to better match speech and text representations. Experiments on multiple benchmarks demonstrate that this approach yields competitive results against strong baselines, highlighting the need to explicitly address structural differences between speech and text in SLM training.
arXiv:2607. 22304v1 Announce Type: new Abstract: Synthetic data augmentation in speech is common practice for linguistic tasks like ASR, but has seen far less work for paralinguistic ones, especially clinical tasks where labelled data is expensive and some patient groups are underrepresented.
The paper proposes a linguistically structured multi‑task learning framework for recognizing non‑canonical phonemes by decomposing phoneme prediction into articulatory feature dimensions such as manner, place, and voicing. A hierarchical architecture with task‑specific heads and a cross‑attention fusion module is combined with semi‑supervised Momentum Pseudo‑Labeling and a cascaded training strategy that gradually introduces articulatory tasks. Experiments on the L2‑ARCTIC dataset demonstrate significant improvements over baseline models and produce interpretable error patterns aligned with phonological feature structure.