arXiv:2609.18533v1 Announce Type: new
Abstract: Automatic speech recognition (ASR) systems exhibit unequal error rates across speaker groups, motivating interventions on their internal representation...
By Nicolas Bourrel, Abderrahmane Issam, Gerasimos Spanakis
Automatic speech recognition (ASR) systems exhibit unequal error rates across speaker groups, motivating interventions on their internal representations. We ask whether speaker-linked attributes that...
arXiv:2609.38106v1 Announce Type: cross
Abstract: Speech-LLMs are expensive to run, making compression important for real-world deployment. However, compressed models are usually selected using aggre...
By Ganesh Pavan Kartikeya Bharadwaj Kolluri, Michael Kampouridis, Ravi Shekhar
arXiv:2609.38976v1 Announce Type: new
Abstract: Demographic fairness gaps in automatic speech recognition are almost always reported from a single training run. We fine-tune the Q-former projector an...
By Srishti Ginjala, Eric Fosler-Lussier, Srinivasan Parthasarathy
arXiv:2609.27382v1 Announce Type: cross
Abstract: Speech technology penalizes some voices: recognition errs nearly twice as often for Black speakers, and accuracy declines for second-language accents...
By Kian Shamsaie, Iman Modarressi
arXiv:2607. 12468v1 Announce Type: cross Abstract: We describe our submission to Task 1 of the 2nd MLCSLM Challenge: a cascaded diarization-then-recognition system that combines DiariZen-Large-s80 (WavLM-Large) segmentation, CAM++ embedding-based two-speaker clustering, and a LoRA-adapted omniASR LLM 7B v2 recognizer, with no oracle segmentation or speaker labels at test time.
By Shuming Fang, Shuifei Zeng
AfriSwitch is a 61.36‑hour, human‑transcribed benchmark of in‑the‑wild code‑switched speech covering 16 African languages and varieties, annotated with switch‑level English span tags, per‑utterance Code‑Mixing Index (CMI), and switch‑point counts. The corpus reveals that code‑switching behaviour varies widely across languages, with no single metric fully capturing how code‑switched a language is. Benchmarking five open and commercial multilingual ASR systems in a zero‑shot setting shows high word error rates, with the best system averaging 35.93% WER and none dropping below 24% on any language, indicating that Africa‑targeted training rather than model scale or nominal language coverage best predicts performance.
By Gabrial Zencha Ashungafac, Busayo Awobade, Tobi Olatunji
Speech technology penalizes some voices: recognition errs nearly twice as often for Black speakers, and accuracy declines for second-language accents and older speakers. We introduce TRIAD, an audit g...
arXiv:2608.30853v1 Announce Type: cross
Abstract: While automatic speech recognition (ASR) models have achieved remarkable improvements in recent years, performance disparities persist across differe...
By Ting-Hui Cheng, Line Katrine Harder Clemmensen, Sneha Das
Demographic fairness gaps in automatic speech recognition are almost always reported from a single training run. We fine-tune the Q-former projector and LoRA adapters of a speech LLM at five audio com...
arXiv:2610.08063v1 Announce Type: cross
Abstract: This paper presents the HINTT system submitted to the 2nd Challenge and Workshop on Multilingual Conversational Speech Language Model (MLC-SLM). We a...
By Takanori Ashihara, Kohei Matsuura, Masato Mimura
arXiv:2608.22196v1 Announce Type: cross
Abstract: While cascaded multi-talker ASR (MT-ASR) leverages state-of-the-art foundation models, its performance is often capped by speaker leakage during sepa...
By Hermann Yepdjio Nkouanga, Minwei Luo, Maggie Wigness, Suresh Singh