arXiv Computation and Language
Sep 25

Temporal Taxation Compounds Under Post-Training Compression of Whisper Models

The paper investigates how post‑training compression techniques—such as pruning, quantization, and distillation—affect demographic fairness in Whisper speech‑recognition models. It finds that pruning and INT4 quantization significantly widen word‑error‑rate gaps between demographic groups, especially for Black/AA and Asian speakers, while distillation tends to reduce these gaps. The study introduces a temporal‑taxation metric to quantify the increased correction effort required for marginalized speakers after compression.

By Srishti Ginjala, Eric Fosler-Lussier, Christopher W. Myers, Srinivasan Parthasarathy