Blazingly fast whisper transcriptions with Inference Endpoints
Related stories
Context-Aware Interleaved Batching for WhisperX
Context-Aware Interleaved Batching for WhisperX proposes a new batching strategy that combines WhisperX’s fast intra-audio batching with Whisper’s sequential context retention. By using voice activity detection (VAD) to define segment boundaries, the algorithm stabilizes Whisper’s text conditioning and preserves continuous historical context across batched audio segments. Experiments on long‑form audio benchmarks show reduced Word Error Rate and improved proper noun transcription while maintaining high‑throughput inference speeds.
MURMUR: An Efficient Inference System for Long-Form ASR
arXiv:2606. 01483v1 Announce Type: cross Abstract: Long-form automatic speech recognition (ASR) requires both high accuracy and low latency, but existing systems force a trade-off between the two.
TTS Arena: Benchmarking Text-to-Speech Models in the Wild
BanglaTurn: A Benchmark and Whisper-Based Model for End-of-Turn Detection in Bangla Speech
BanglaTurn is a new corpus of 35,374 Bangla podcast speech samples, each 3 to 15 seconds long, labeled for end‑of‑turn detection through speaker diarization, an LLM pass, and human verification. A Whisper‑based model with task‑specific classification heads achieves 84.33 % accuracy on a balanced test set, outperforming the Smart‑Turn v3 baseline (69.28 %) and reducing the false‑negative rate from 51.57 % to 7.55 %, though with a higher false‑positive rate. The study also details the contributions of encoder‑layer fine‑tuning, multi‑scale pooling, INT8 quantization, and reports inference latency of 165–191 ms on CPU.
RelayS2S: A Dual-Path Speculative Generation for Real-Time Dialogue
RelayS2S is a hybrid real‑time dialogue system that runs a fast duplex speech‑to‑speech path and a slow ASR‑to‑LLM path in parallel. The fast path speculatively drafts a short response prefix and streams it to TTS, while the slow path generates a higher‑quality continuation conditioned on that prefix. A lightweight verifier decides whether to commit the prefix or fall back to the cascaded pipeline, achieving much lower latency (81 ms P90 first‑chunk) while preserving 99% of the cascaded pipeline’s textual quality.
Deterministic Prompting for Speaker-Stable Low-Resource Greek TTS
Modern TTS systems approach human quality for high-resource languages but degrade when clean speech data is scarce. Modern Greek exemplifies this, lacking the curated corpora behind state-of-the-art s...
Introducing Whisper
Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition
arXiv:2608. 10836v1 Announce Type: cross Abstract: The signal ambiguity of whispered speech drives ASR systems toward two opposing failure modes: failing to capture whispered speech or hallucinatory transcription of noise.
X2Streaming-ASR: wait when uncertain, emit when ready for streaming ASR
X2Streaming-ASR introduces a method for streaming automatic speech recognition that separates the decision of when to commit a transcript from what to commit. The approach uses a three‑stage training process: first establishing streaming capability, then warm‑starting a commit policy with automatically probed trajectories, and finally refining the policy with character‑level, segment‑assigned group‑relative rewards for accuracy and latency. On AISHELL‑1/2/3 and WenetSpeech datasets, the system achieves mean character‑level commit latencies of 27–84 ms, far lower than baseline systems, while also attaining the best streaming character error rates on AISHELL‑1 and AISHELL‑3.
Deterministic Prompting for Speaker-Stable Low-Resource Greek TTS
arXiv:2609.10022v1 Announce Type: cross Abstract: Modern TTS systems approach human quality for high-resource languages but degrade when clean speech data is scarce. Modern Greek exemplifies this, la...
Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music
arXiv:2412. 11449v2 Announce Type: replace-cross Abstract: We propose WHISPER-GPT: A generative large language model (LLM) for speech and music that allows us to work with continuous audio representations and discrete tokens simultaneously as part of a single architecture.