Hugging Face Blog

Blazingly fast whisper transcriptions with Inference Endpoints

arXiv Computation and Language
Sep 1

Context-Aware Interleaved Batching for WhisperX

Context-Aware Interleaved Batching for WhisperX proposes a new batching strategy that combines WhisperX’s fast intra-audio batching with Whisper’s sequential context retention. By using voice activity detection (VAD) to define segment boundaries, the algorithm stabilizes Whisper’s text conditioning and preserves continuous historical context across batched audio segments. Experiments on long‑form audio benchmarks show reduced Word Error Rate and improved proper noun transcription while maintaining high‑throughput inference speeds.

By Carlos Bain, Max Bain
arXiv Computation and Language
Sep 25

BanglaTurn: A Benchmark and Whisper-Based Model for End-of-Turn Detection in Bangla Speech

BanglaTurn is a new corpus of 35,374 Bangla podcast speech samples, each 3 to 15 seconds long, labeled for end‑of‑turn detection through speaker diarization, an LLM pass, and human verification. A Whisper‑based model with task‑specific classification heads achieves 84.33 % accuracy on a balanced test set, outperforming the Smart‑Turn v3 baseline (69.28 %) and reducing the false‑negative rate from 51.57 % to 7.55 %, though with a higher false‑positive rate. The study also details the contributions of encoder‑layer fine‑tuning, multi‑scale pooling, INT8 quantization, and reports inference latency of 165–191 ms on CPU.

By Mizbaul Haque Maruf
arXiv AI
Sep 11

RelayS2S: A Dual-Path Speculative Generation for Real-Time Dialogue

RelayS2S is a hybrid real‑time dialogue system that runs a fast duplex speech‑to‑speech path and a slow ASR‑to‑LLM path in parallel. The fast path speculatively drafts a short response prefix and streams it to TTS, while the slow path generates a higher‑quality continuation conditioned on that prefix. A lightweight verifier decides whether to commit the prefix or fall back to the cascaded pipeline, achieving much lower latency (81 ms P90 first‑chunk) while preserving 99% of the cascaded pipeline’s textual quality.

By Long Mai, Junli Liang
arXiv AI
Sep 10

X2Streaming-ASR: wait when uncertain, emit when ready for streaming ASR

X2Streaming-ASR introduces a method for streaming automatic speech recognition that separates the decision of when to commit a transcript from what to commit. The approach uses a three‑stage training process: first establishing streaming capability, then warm‑starting a commit policy with automatically probed trajectories, and finally refining the policy with character‑level, segment‑assigned group‑relative rewards for accuracy and latency. On AISHELL‑1/2/3 and WenetSpeech datasets, the system achieves mean character‑level commit latencies of 27–84 ms, far lower than baseline systems, while also attaining the best streaming character error rates on AISHELL‑1 and AISHELL‑3.

By Zhiwei Lin, Kaiqi Fu, Rime Wen, Zehan Liu, Shawn Qin, Roy Gan, Hao Wang, Qian Wang