Blazingly fast whisper transcriptions with Inference Endpoints
Related stories
MURMUR: An Efficient Inference System for Long-Form ASR
arXiv:2606. 01483v1 Announce Type: cross Abstract: Long-form automatic speech recognition (ASR) requires both high accuracy and low latency, but existing systems force a trade-off between the two.
TTS Arena: Benchmarking Text-to-Speech Models in the Wild
Introducing Whisper
Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition
arXiv:2608. 10836v1 Announce Type: cross Abstract: The signal ambiguity of whispered speech drives ASR systems toward two opposing failure modes: failing to capture whispered speech or hallucinatory transcription of noise.
Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music
arXiv:2412. 11449v2 Announce Type: replace-cross Abstract: We propose WHISPER-GPT: A generative large language model (LLM) for speech and music that allows us to work with continuous audio representations and discrete tokens simultaneously as part of a single architecture.
MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation
arXiv:2607. 26698v1 Announce Type: cross Abstract: Cover song generation (CSG) should preserve the melodic and linguistic content of a reference song while recreating the remaining musical components.
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation
arXiv:2608. 11590v1 Announce Type: cross Abstract: Human voice generation has made rapid progress in speech generation, singing voice generation, voice cloning, and voice editing.
NaturalFlow: Reducing Disruptive Pauses for Natural Speech Flow in Simultaneous Speech-to-Speech Translation
arXiv:2606. 13121v1 Announce Type: cross Abstract: Simultaneous speech-to-speech translation aims to enable near-real-time communication by minimizing latency, offering a compelling, real-time alternative to the high latency of consecutive translation.
Introducing the Realtime API
Developers can now build fast speech-to-speech experiences into their applications
ZONOS2 Technical Report
We present ZONOS2 8B, our latest TTS model, which achieves state-of-the-art naturalness, prosody, and voice cloning fidelity. We improve upon Zonos-v0.
Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS
arXiv:2605. 30748v2 Announce Type: replace-cross Abstract: We present Chatterbox-Flash, a zero-shot text-to-speech model obtained by fine-tuning a pretrained autoregressive TTS decoder into a block-diffusion decoder, enabling parallel token generation within each block while retaining block-by-block streaming.