Boosting Wav2Vec2 with n-grams in 🤗 Transformers
Related stories
Fine-Tune XLSR-Wav2Vec2 for low-resource ASR with 🤗 Transformers
Fine-Tune Wav2Vec2 for English ASR in Hugging Face with 🤗 Transformers
Fine-Tune W2V2-Bert for low-resource ASR with 🤗 Transformers
Training a language model with 🤗 Transformers using TensorFlow and TPUs
Overview of natively supported quantization schemes in 🤗 Transformers
TontaubeV1: Streaming Text-to-Speech with Hierarchical Codec Modeling and Bounded Context
TontaubeV1 is a streaming text‑to‑speech model that maintains natural prosody while running on a single consumer GPU. It encodes speech with a hierarchical DualCodec representation at 12.5 Hz, separating a semantic stream from successive acoustic refinements, and uses Qwen3‑derived transformers to predict the semantic stream and add refinements. The model supports up to one minute of reference audio for voice conditioning, streams with a 200 ms latency to first audio, and achieves real‑time factors of 0.08 (single input) and 0.02 (eight concurrent inputs).
Kraken: LLM-based Speech-to-Speech Translation via Low-bitrate VQ and Dual-path Source Conditioning
arXiv:2609.13045v1 Announce Type: new Abstract: Speech-to-speech translation (S2ST) has advanced significantly with speech LLMs, offering the potential for joint optimization and preserving non-lingu...
Phonetic Error Analysis of Raw Waveform Acoustic Models
arXiv:2606. 07030v1 Announce Type: cross Abstract: We analyse error patterns of raw waveform acoustic models on TIMIT phone recognition beyond the overall phone error rate (PER).
HybridCodec: Modeling Discrete and Continuous Representations for Efficient Speech Language Models
arXiv:2606. 27627v1 Announce Type: cross Abstract: Discrete audio representations have become increasingly popular for building multimodal text-audio systems and integrating audio capabilities into Large Language Models (LLMs).
TontaubeV1: Streaming Text-to-Speech with Hierarchical Codec Modeling and Bounded Context
TontaubeV1 is a streaming text‑to‑speech model that preserves natural prosody while running on a single consumer GPU. It encodes speech with a hierarchical DualCodec representation at 12.5 Hz, separating a semantic stream from successive acoustic refinements, and uses Qwen3‑derived transformers to predict the semantic stream, utterance duration, and acoustic refinements. The system supports up to one minute of reference audio for voice conditioning, streams with a 200 ms latency to first audio, and achieves real‑time factors of 0.08 (single input) and 0.02 (eight concurrent inputs).