arXiv Machine Learning By Fiza Husain, Ankit Pandey, Yash Singh

Beyond WER: Entity and Disfluency Recall in Accented Conversational ASR

Read the original on arXiv Machine Learning →

The paper introduces a three‑stage pipeline to improve accented conversational ASR for speakers from India, Indonesia, and Latin America. It uses heuristic SQL filters to curate entity‑rich training data, regional LoRA adapters fine‑tuned on Qwen2.5‑Omni‑3B to generate both verbatim and corrected transcripts, and a six‑category error taxonomy validated by an LLM judge. The approach raises entity recall to 80‑85% and filler recall to 76‑86%, while keeping WER low (6‑10%) and outperforming Whisper and a commercial ASR on entity recall.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 9

Subtitle-Aligned Fine-Tuning of Whisper for Swiss German ASR: Benchmark Contamination, Convention Mismatch, and an Honest Baseline at 25.6% WER (13.8% cWER)

arXiv:2606. 07608v1 Announce Type: cross Abstract: We present a systematic study of fine-tuning OpenAI's Whisper large-v3 for Swiss German ASR, using 1,367 hours of broadcast speech paired with Standard German subtitles as weak supervision.

By Felix Akeret
arXiv Computation and Language
Aug 28

Scaling phoneme-based TTS augmentation for ASR: A unified pipeline and controlled study

The paper introduces a unified phoneme‑based TTS‑to‑ASR augmentation pipeline that uses a multilingual TTS model with language‑ID conditioning and incorporates grapheme‑to‑phoneme conversion, reference‑speech filtering, and candidate‑text selection. It proposes phoneme‑frequency‑guided selection (PFGS) to rank sentences based on phoneme frequencies from real ASR labels, and demonstrates that random augmentation and PFGS both improve ASR performance across Arabic, French, Italian, and Portuguese test sets, with PFGS yielding up to a 19.3% relative WER reduction. The study also shows that filtering reference speech can further lower WER by up to 0.59 points on certain datasets.

By Zhen Wang, TianRui Wu, RongQi Han, Hao Wu, Wei Liang