Boosting Wav2Vec2 with n-grams in π€ Transformers
Related stories
Fine-Tune XLSR-Wav2Vec2 for low-resource ASR with π€ Transformers
Fine-Tune Wav2Vec2 for English ASR in Hugging Face with π€ Transformers
Fine-Tune W2V2-Bert for low-resource ASR with π€ Transformers
Training a language model with π€Β Transformers using TensorFlow and TPUs
Overview of natively supported quantization schemes in π€ Transformers
Phonetic Error Analysis of Raw Waveform Acoustic Models
arXiv:2606. 07030v1 Announce Type: cross Abstract: We analyse error patterns of raw waveform acoustic models on TIMIT phone recognition beyond the overall phone error rate (PER).
HybridCodec: Modeling Discrete and Continuous Representations for Efficient Speech Language Models
arXiv:2606. 27627v1 Announce Type: cross Abstract: Discrete audio representations have become increasingly popular for building multimodal text-audio systems and integrating audio capabilities into Large Language Models (LLMs).
Fine-Tune Whisper For Multilingual ASR with π€ Transformers
UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations
arXiv:2604. 14606v2 Announce Type: cross Abstract: Universal speech enhancement (USE) aims to restore speech signals from diverse distortions across multiple sampling rates.
How to train a new language model from scratch using Transformers and Tokenizers
ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models
arXiv:2607. 09134v1 Announce Type: cross Abstract: Representation alignment (REPA) has been investigated to accelerate diffusion training, but we observe that regularizing intermediate representations in diffusion Transformers (DiT) may implicitly entangle latents and limit generative capacity.