Fine-Tune XLSR-Wav2Vec2 for low-resource ASR with 🤗 Transformers
Related stories
Fine-Tune Wav2Vec2 for English ASR in Hugging Face with 🤗 Transformers
Boosting Wav2Vec2 with n-grams in 🤗 Transformers
Making automatic speech recognition work on large files with Wav2Vec2 in 🤗 Transformers
Fine-Tune Whisper For Multilingual ASR with 🤗 Transformers
Fine-Tune MMS Adapter Models for low-resource ASR
ZipCodec: Ultra-Low-Frame-Rate Streaming Speech Coding
ZipCodec is a streaming neural speech codec that operates at an ultra‑low frame rate of 6.25 Hz and a bitrate of 0.80 kbps, achieving a theoretical latency of 160 ms. It leverages large‑scale WavLM distillation, a redesigned transformer architecture, a scalar spherical quantizer, and a latency‑aware streaming decoder to preserve reconstruction quality while reducing frame rate. Experiments demonstrate that ZipCodec outperforms existing streaming codecs at comparable bitrates in both reconstruction and downstream tasks, and it can run real‑time single‑stream inference on a consumer‑grade CPU despite having 842 M parameters.
ZipCodec: Ultra-Low-Frame-Rate Streaming Speech Coding
ZipCodec is a streaming neural speech codec that operates at an ultra‑low frame rate of 6.25 Hz and a bitrate of 0.80 kbps, achieving a theoretical latency of 160 ms. It leverages large‑scale WavLM distillation, a redesigned transformer architecture, a scalar spherical quantizer, and a latency‑aware streaming decoder. Experiments demonstrate that ZipCodec outperforms existing streaming codecs at comparable bitrates in both reconstruction quality and downstream tasks, while remaining real‑time on a consumer‑grade CPU despite its 842 M parameters.
Contrastive Regularization for Accent-Robust ASR
arXiv:2605. 03297v2 Announce Type: replace-cross Abstract: ASR systems based on self-supervised acoustic pretraining and CTC fine-tuning achieve strong performance on native speech but remain sensitive to accent variability.
TontaubeV1: Streaming Text-to-Speech with Hierarchical Codec Modeling and Bounded Context
TontaubeV1 is a streaming text‑to‑speech model that maintains natural prosody while running on a single consumer GPU. It encodes speech with a hierarchical DualCodec representation at 12.5 Hz, separating a semantic stream from successive acoustic refinements, and uses Qwen3‑derived transformers to predict the semantic stream and add refinements. The model supports up to one minute of reference audio for voice conditioning, streams with a 200 ms latency to first audio, and achieves real‑time factors of 0.08 (single input) and 0.02 (eight concurrent inputs).
Benchmarking_Fast_Domain_Adaptation_for_Unsupervised_Speech_Units
The paper introduces ABX-Accent, a benchmark built on the AESRC dataset that evaluates how well representation learning models adapt to 10 different English accents with less than 10 hours of unlabeled data per accent. It adapts the Zero Resources Challenge ABX metrics for each accent and demonstrates a baseline using adaptive domain normalization to fine‑tune a Contrastive Predictive Coding model, achieving a 23.6% relative improvement on across‑speaker ABX scores compared to non‑adapted models. The dataset and evaluation metrics will be released publicly after the paper is accepted.
TontaubeV1: Streaming Text-to-Speech with Hierarchical Codec Modeling and Bounded Context
TontaubeV1 is a streaming text‑to‑speech model that preserves natural prosody while running on a single consumer GPU. It encodes speech with a hierarchical DualCodec representation at 12.5 Hz, separating a semantic stream from successive acoustic refinements, and uses Qwen3‑derived transformers to predict the semantic stream, utterance duration, and acoustic refinements. The system supports up to one minute of reference audio for voice conditioning, streams with a 200 ms latency to first audio, and achieves real‑time factors of 0.08 (single input) and 0.02 (eight concurrent inputs).