arXiv Computation and Language

Fairness Beyond a Single Run: Training-Seed Variability in Speech LLM Adaptation

arXiv Computation and Language
Sep 25

Temporal Taxation Compounds Under Post-Training Compression of Whisper Models

The paper investigates how post‑training compression techniques—such as pruning, quantization, and distillation—affect demographic fairness in Whisper speech‑recognition models. It finds that pruning and INT4 quantization significantly widen word‑error‑rate gaps between demographic groups, especially for Black/AA and Asian speakers, while distillation tends to reduce these gaps. The study introduces a temporal‑taxation metric to quantify the increased correction effort required for marginalized speakers after compression.

By Srishti Ginjala, Eric Fosler-Lussier, Christopher W. Myers, Srinivasan Parthasarathy
arXiv Computation and Language
Aug 31

SURE-Challenge: Evaluating Speech Evidence Before Speech-LLM Generation

The paper introduces the Speech-Unsupported Rejection Evaluation Challenge (SURE‑Challenge), a benchmark designed to test whether speech‑LLMs should accept or reject audio inputs before generating answers. Using LibriSpeech‑derived transcriptions paired with first‑word question answering, the authors evaluate various noise and silence conditions, and compare a simple energy‑plus‑Whisper‑score rule against a Qwen2‑Audio front‑end. On a 474‑row test set, the rule rejects 196 of 204 unsupported inputs while preserving accuracy on supported data, revealing a pre‑generation error mode that answer‑only scoring misses.

By Mengzhe Geng