Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The paper introduces a method to prune six layers from the encoder of OpenAI’s Whisper ASR model, reducing the encoder stack by 18.5% without requiring custom inference code. Layers are selected based on their minimal impact on Word Error Rate when removed. After pruning, the model’s WER rises from 18.2% to 21.9%, but distillation with unlabeled monolingual speech data lowers it to 20.1%. "whyItMatters":"The approach offers a straightforward way to accelerate Whisper inference by simplifying the encoder while maintaining acceptable accuracy, and the released code and model enable immediate adoption by the community."
arXiv:2511.08092v2 Announce Type: replace-cross Abstract: We challenge the conventional view of neural network pruning as solely a compression technique, demonstrating that one-shot magnitude pruning...
X-AuT is a progressive framework for compressing the audio encoder of speech large language models. It selects layer combinations via short behavioral probes and restores performance through representation alignment, cross‑scale distillation, scheduled student‑policy supervision, and LoRA finetuning, while keeping the language‑model backbone frozen. On ten Chinese–English benchmarks, reducing Qwen3‑ASR‑0.6B’s encoder from 18 to 16 layers lowers macro‑average error from 5.61% to 5.27%, and a 14‑layer model achieves 5.75% error with 20.7% fewer parameters.
arXiv:2603. 05121v2 Announce Type: replace-cross Abstract: Speech Large Language Models route speech encoder representations into an LLM decoder that typically accounts for over 90% of total parameters.
arXiv:2508. 07048v2 Announce Type: replace-cross Abstract: Autoregressive (AR) encoder-decoder models dominate high-quality multilingual ASR, but their left-to-right decoders make inference latency scale with transcript length.
arXiv:2608. 06630v1 Announce Type: new Abstract: Pruning reduces the inference cost of large language models, but existing criteria primarily preserve large activations or reconstruct layer outputs.