Fine-Tune MMS Adapter Models for low-resource ASR
Related stories
Fine-Tune W2V2-Bert for low-resource ASR with 🤗 Transformers
Fine-Tune Whisper For Multilingual ASR with 🤗 Transformers
Introducing the FFASR Leaderboard: Benchmarking ASR in the Real World
Fine-tuning LLMs to 1.58bit: extreme quantization made easy
Scaling Audio Models Efficiently: Joint Optimization of Scale, Resolution, Adaptation, Precision, and Sparsity
The paper introduces a compression framework for the Whisper automatic speech recognition model that jointly optimizes six deployment dimensions—model size, temporal resolution, encoder token stride, low‑rank adaptation capacity, weight precision, and sparsity pattern—using NSGA‑III. The optimization targets three objectives: word error rate, inference FLOPs, and memory footprint. Evaluating 1,680 configurations, the study identifies compression combinations that outperform single‑axis scaling and notes that 1:4 structured sparsity cannot maintain acceptable accuracy within the tested budgets.
Diving into Kronecker Adapters: Component Design Matters
arXiv:2602. 01267v3 Announce Type: replace Abstract: Kronecker adapters have emerged as a promising approach for fine-tuning large-scale models, enabling high-rank updates through tunable component structures.
Data Scale, Not Latency, Shapes Cross-Lingual Encoder Transfer in Streaming ASR
arXiv:2606. 24169v1 Announce Type: new Abstract: Adapting a streaming speech recognition model to a new language requires choosing between two plausible warm starts: a multilingual (ML) encoder or an English-only (EN) encoder.
ZipCodec: Ultra-Low-Frame-Rate Streaming Speech Coding
ZipCodec is a streaming neural speech codec that operates at an ultra‑low frame rate of 6.25 Hz and a bitrate of 0.80 kbps, achieving a theoretical latency of 160 ms. It leverages large‑scale WavLM distillation, a redesigned transformer architecture, a scalar spherical quantizer, and a latency‑aware streaming decoder to preserve reconstruction quality while reducing frame rate. Experiments demonstrate that ZipCodec outperforms existing streaming codecs at comparable bitrates in both reconstruction and downstream tasks, and it can run real‑time single‑stream inference on a consumer‑grade CPU despite having 842 M parameters.
MURMUR: An Efficient Inference System for Long-Form ASR
arXiv:2606. 01483v1 Announce Type: cross Abstract: Long-form automatic speech recognition (ASR) requires both high accuracy and low latency, but existing systems force a trade-off between the two.
Sequential Adapter Stacking for Cross-Lingual Low-Resource ASR
arXiv:2609.15758v1 Announce Type: new Abstract: Extending large-scale multilingual automatic speech recognition (ASR) models to low-resource languages remains challenging. Model performance is skewed...
Interface-Aware KV Cache Quantization for Dense On-Chip NVM in Long-Context LLM Decoding
arXiv:2609.05764v1 Announce Type: cross Abstract: The key-value (KV) cache is the dominant memory bottleneck in long-context large language model (LLM) decoding: every step reads it entirely, so deco...