Voxtral transcribes at the speed of sound.
Related stories
Speaking of Voxtral
Voxtral TTS: A frontier, open-weights text-to-speech model that’s fast, instantly adaptable, and produces lifelike speech for voice agents.
AudioLDM 2, but faster ⚡️
Spectro-Temporal Interference Confounds Phase Encoding in Spatial Audio Foundation Models
arXiv:2606. 14820v1 Announce Type: cross Abstract: Recent spatial self supervised audio models achieve high performance on localization tasks, raising questions about their encoding of microsecond interaural phase fine structures.
Blazingly fast whisper transcriptions with Inference Endpoints
VoxZip: Semantic-Anchored Temporal KV Cache Compression for Long-Context Audio Inference
arXiv:2608. 08569v1 Announce Type: new Abstract: Recent advancements in Speech Large Language Models have demonstrated remarkable capabilities in understanding complex audio tasks.
Probing Low Frame Rate Degradation in Neural Audio Codecs
arXiv:2606. 16969v1 Announce Type: cross Abstract: Low frame rates in neural audio codecs are attractive for autoregressive speech synthesis, where the generation cost scales linearly with the sequence length.
Improved Gemini audio models for powerful voice experiences
Fast When, Careful Who: Dual-Process Multiparty Turn-Taking with Diffusion Augmentation
arXiv:2606. 16568v1 Announce Type: cross Abstract: Reliable turn-taking is essential for spoken dialogue systems.
Whisper Hallucination Detection and Mitigation via Hidden Representation Steering and Sparse AutoEncoders
arXiv:2606. 07473v1 Announce Type: cross Abstract: Whisper, a widely adopted ASR model, is known to suffer from hallucinations - coherent transcriptions generated for non-speech audio entirely disconnected from the input.
VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion
arXiv:2607. 11706v1 Announce Type: cross Abstract: Modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represented in many legacy spoofing benchmarks.
Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval
Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2. 0 audio embeddings.