arXiv:2609.38976v1 Announce Type: new
Abstract: Demographic fairness gaps in automatic speech recognition are almost always reported from a single training run. We fine-tune the Q-former projector an...
By Srishti Ginjala, Eric Fosler-Lussier, Srinivasan Parthasarathy
arXiv:2609.38106v1 Announce Type: cross
Abstract: Speech-LLMs are expensive to run, making compression important for real-world deployment. However, compressed models are usually selected using aggre...
By Ganesh Pavan Kartikeya Bharadwaj Kolluri, Michael Kampouridis, Ravi Shekhar
arXiv:2609.18533v1 Announce Type: new
Abstract: Automatic speech recognition (ASR) systems exhibit unequal error rates across speaker groups, motivating interventions on their internal representation...
By Nicolas Bourrel, Abderrahmane Issam, Gerasimos Spanakis
Automatic speech recognition (ASR) systems exhibit unequal error rates across speaker groups, motivating interventions on their internal representations. We ask whether speaker-linked attributes that...
The paper investigates how post‑training compression techniques—such as pruning, quantization, and distillation—affect demographic fairness in Whisper speech‑recognition models. It finds that pruning and INT4 quantization significantly widen word‑error‑rate gaps between demographic groups, especially for Black/AA and Asian speakers, while distillation tends to reduce these gaps. The study introduces a temporal‑taxation metric to quantify the increased correction effort required for marginalized speakers after compression.
By Srishti Ginjala, Eric Fosler-Lussier, Christopher W. Myers, Srinivasan Parthasarathy
The paper introduces the Speech-Unsupported Rejection Evaluation Challenge (SURE‑Challenge), a benchmark designed to test whether speech‑LLMs should accept or reject audio inputs before generating answers. Using LibriSpeech‑derived transcriptions paired with first‑word question answering, the authors evaluate various noise and silence conditions, and compare a simple energy‑plus‑Whisper‑score rule against a Qwen2‑Audio front‑end. On a 474‑row test set, the rule rejects 196 of 204 unsupported inputs while preserving accuracy on supported data, revealing a pre‑generation error mode that answer‑only scoring misses.
By Mengzhe Geng
arXiv:2609.27382v1 Announce Type: cross
Abstract: Speech technology penalizes some voices: recognition errs nearly twice as often for Black speakers, and accuracy declines for second-language accents...
By Kian Shamsaie, Iman Modarressi
Speech technology penalizes some voices: recognition errs nearly twice as often for Black speakers, and accuracy declines for second-language accents and older speakers. We introduce TRIAD, an audit g...
arXiv:2608.27783v3 Announce Type: replace-cross
Abstract: Speech language models (speech LLMs) can generate plausible outputs from audio that contains no usable speech evidence. We study this failure...
By Mengzhe Geng
arXiv:2606. 10911v1 Announce Type: cross Abstract: Claims about the robustness and fairness of deepfake speech detectors are only as credible as the datasets used to train and evaluate those systems.
By Vojt\v{e}ch Stan\v{e}k, Eva Trnovsk\'a, Kamil Malinka, Anton Firc
arXiv:2609.23825v1 Announce Type: new
Abstract: We present a comprehensive benchmark of Federated Learning (FL) for multilingual Automatic Speech Recognition (ASR), evaluating four Speech-LLM archite...
By Jordi Luque, Aleix Sant, Fernando L\'opez
arXiv:2609.36500v1 Announce Type: cross
Abstract: Speaker verification systems encounter combinations of noise, channel distortion, and changes in speech. Evaluating each condition separately does no...
By Kamel Kamel, Hridoy Sankar Dutta, Keshav Sood, Sunil Aryal