arXiv AI

Schr\"odinger Bridge Mamba for One-Step Speech Enhancement

arXiv:2510. 16834v3 Announce Type: replace-cross Abstract: We present Schr\"odinger Bridge Mamba (SBM), a novel model for efficient speech enhancement by integrating the Schr\"odinger Bridge (SB) training paradigm and the Mamba architecture.

arXiv AI
Jun 9

Speech Enhancement Based on Drifting Models

arXiv:2604. 24199v4 Announce Type: replace-cross Abstract: We propose Speech Enhancement based on Drifting Models (DriftSE), a novel generative framework that formulates denoising as an equilibrium problem.

By Liang Xu, Diego Caviedes-Nozal, W. Bastiaan Kleijn, Longfei Felix Yan, Rasmus Kongsgaard Olsson
arXiv AI
Jun 30

How to Leverage Synthetic Speech for LLM-Based ASR Systems?

arXiv:2606. 29031v1 Announce Type: cross Abstract: In regulated domains such as banking and healthcare, where privacy constraints make real speech costly to collect and retain, synthetic speech from modern text-to-speech (TTS) is an appealing alternative for training automatic speech recognition (ASR) without exposing sensitive customer recordings.

By Yanis Labrak, Dairazalia Sanchez-Cortes, Sergio Burdisso, S\'everin Baroudi, Shashi Kumar, Esa\'u Villatoro-Tello, Srikanth Madikeri, Manjunath K E, Old\v{r}ich Plchot, Kadri Hacio\u{g}lu, Petr Motlicek, Andreas Stolcke
arXiv Machine Learning
Sep 25

Transcript-Supervised Post-Training of Generative Speech Enhancement on Real Recordings via Reinforce Adjoint Matching

The paper adapts Reinforce Adjoint Matching (RAM) to generative speech enhancement, allowing a pretrained model to be post‑trained on real recordings using weak supervision such as text transcripts. RAM shifts the model’s conditional distribution toward higher‑reward outputs by generating enhanced speech on‑policy, evaluating each output with a potentially non‑differentiable reward, and analytically re‑noising the endpoint to create inputs for a reward‑guided regression objective. Experiments on real CHiME‑4 recordings show a 5.08‑percentage‑point reduction in word error rate compared to the pretrained FlowSE model, while maintaining all reported non‑intrusive speech quality metrics and receiving no significant preference in a subjective listening test.

By Julius Richter, Christoph Boeddeker, Yoshiki Masuyama, Kohei Saijo, Dominik Klement, Gordon Wichern, Jonathan Le Roux
arXiv Computation and Language
Sep 17

RT-SEMamba: Real-Time Speech Enhancement Mamba via Progressive Knowledge Distillation

RT-SEMamba is a fully causal speech enhancement model that uses causal time‑frequency Mamba blocks instead of Transformer‑based architectures, allowing efficient long‑form inference with a fixed‑size recurrent state. The authors introduce a progressive knowledge distillation strategy that compresses an 8‑layer teacher into a single‑layer student by jointly distilling spectral outputs and intermediate representations. On the Voicebank‑DEMAND benchmark, the 8‑layer model achieves 3.32 PESQ under a 25 ms latency constraint, while the distilled 1‑layer student improves from 3.06 to 3.18 PESQ, maintains the same steady‑state real‑time factor, and runs 2.64× faster than the teacher.

By Rong Chao, Sung-Feng Huang, Moreno La Quatra, Sabato Marco Siniscalchi, Wen-Huang Cheng, Szu-Wei Fu, Yu Tsao