arXiv:2609.21663v1 Announce Type: new
Abstract: Word Error Rate (WER), the most commonly used metric for Automatic Speech Recognition (ASR), treats every lexical deviation from the reference as equal...
By Hritika Sharma, Thibault Ba\~neras-Roux, Alessandra Pinto, Petr Motlicek, Hyunggu Jung, Esa\'u Villatoro-Tello, Somang Nam
The paper introduces a unified phoneme‑based TTS‑to‑ASR augmentation pipeline that uses a multilingual TTS model with language‑ID conditioning and incorporates grapheme‑to‑phoneme conversion, reference‑speech filtering, and candidate‑text selection. It proposes phoneme‑frequency‑guided selection (PFGS) to rank sentences based on phoneme frequencies from real ASR labels, and demonstrates that random augmentation and PFGS both improve ASR performance across Arabic, French, Italian, and Portuguese test sets, with PFGS yielding up to a 19.3% relative WER reduction. The study also shows that filtering reference speech can further lower WER by up to 0.59 points on certain datasets.
By Zhen Wang, TianRui Wu, RongQi Han, Hao Wu, Wei Liang
The paper examines how multilingual medical adaptation affects the internal representations of Whisper ASR models. By comparing various fine‑tuning strategies—zero‑shot decoding, English‑only, German‑only, two‑stage EN→EN+DE, and direct EN+DE fine‑tuning—it shows that fine‑tuning significantly improves performance, with the best model varying by setting. Layer‑wise encoder analysis reveals that English medical fine‑tuning drives the main representation shift, while multilingual continuation largely preserves the adapted space, and that domain and language signals remain recoverable across layers.
By Souranil Kahali, Rituparna Bose, Abner Hernandez, Tomas Arias-Vergara, Andreas Maier, Ning Ma, Paula Andrea Perez-Toro
The paper compares encoder‑based and generative decoder‑based large language models for evaluating automatic speech recognition (ASR). It examines BERTScore and SemDist across various LLMs, layers, and pooling strategies, finding that both metrics can strongly correlate with human judgments when properly configured. For generative LLMs, the study explores pairwise hypothesis selection via prompting and direct error classification, showing that while encoder‑based metrics remain competitive, generative models excel in hypothesis comparison and enhance interpretability of ASR evaluation.
By Thibault Ba\~neras-Roux, Shashi Kumar, Driss Khalil, Sergio Burdisso, Petr Motlicek, Shiran Liu, Mickael Rouvier, Jane Wottawa, Richard Dufour
This study benchmarks transformer models for Bangla medical named entity recognition (NER), comparing BanglaBERT, multilingual BERT (mBERT), XLM‑RoBERTa, and GPT‑4o mini under zero‑shot and few‑shot prompting. Across a full test set of 3,179 samples, fine‑tuned XLM‑RoBERTa achieves a new state‑of‑the‑art F1‑score of 0.5959, while BanglaBERT lags with 0.4937, suggesting that domain diversity outweighs language specificity. The analysis shows high performance on Medicine and Specialist entities (F1 > 0.83) but lower accuracy on Symptoms (F1 0.4367), and demonstrates that fine‑tuned transformers outperform prompt‑only approaches by a factor of 3.76.
By Rakib Abdullah, Md. Maruful Islam Maruf
The paper examines how multilingual medical adaptation affects the internal representations of Whisper ASR models by performing layer‑wise encoder analysis. It compares several adaptation strategies—zero‑shot decoding, English‑only fine‑tuning, German‑only diagnostic fine‑tuning, two‑stage EN→EN+DE continuation, and direct EN+DE fine‑tuning—across different Whisper sizes, finding that fine‑tuning improves performance but the best model varies by setting. Layer‑wise results show that English medical fine‑tuning drives the main encoder shift, while multilingual continuation largely preserves the adapted representation space, with domain and language information remaining recoverable across layers.
"whyItMatters":"The study provides insight into how multilingual medical adaptation reshapes Whisper’s internal representations, guiding the selection of model sizes and fine‑tuning strategies for improved MedASR performance."