arXiv Machine Learning

Transcript-Supervised Post-Training of Generative Speech Enhancement on Real Recordings via Reinforce Adjoint Matching

The paper adapts Reinforce Adjoint Matching (RAM) to generative speech enhancement, allowing a pretrained model to be post‑trained on real recordings using weak supervision such as text transcripts. RAM shifts the model’s conditional distribution toward higher‑reward outputs by generating enhanced speech on‑policy, evaluating each output with a potentially non‑differentiable reward, and analytically re‑noising the endpoint to create inputs for a reward‑guided regression objective. Experiments on real CHiME‑4 recordings show a 5.08‑percentage‑point reduction in word error rate compared to the pretrained FlowSE model, while maintaining all reported non‑intrusive speech quality metrics and receiving no significant preference in a subjective listening test.

arXiv AI
Jul 16

LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement

arXiv:2603. 13952v3 Announce Type: replace-cross Abstract: In existing Audio-Visual Speech Enhancement (AVSE) methods, objectives such as Scale-Invariant Signal-to-Noise Ratio (SI-SNR) and Mean Squared Error (MSE) are widely used; however, their correlation with perceived speech quality is often suboptimal and provides limited interpretability for optimization.

By Chih-Ning Chen, Jen-Cheng Hou, Hsin-Min Wang, Shao-Yi Chien, Yu Tsao, Fan-Gang Zeng
arXiv AI
Sep 2

Cleaner Speech, Weaker Generalization: Revisiting Pitt-Derived Benchmarks for Alzheimer's Disease Detection

The study examines how speech preprocessing—such as enhancement, sample selection, and demographic balancing—affects Alzheimer’s disease detection models that use the Pitt Corpus. Experiments reveal that while speech‑enhanced datasets boost in‑domain accuracy, they diminish cross‑dataset robustness and introduce class imbalance and prediction shifts, even when training and testing enhancements are matched. Large audio‑language models show similar sensitivity, indicating that cleaner speech does not guarantee better real‑world performance.

By Luqi Sun, Shreeram Suresh Chandra, Lin Zhang, You-Jin Li, Brian MacWhinney, Yu Tsao, Emily Mower Provost, Berrak Sisman
arXiv Machine Learning
Sep 22

Corrective Forcing: Unified Post-Training for Diffusions and Flows in Generative Speech Enhancement

The paper introduces Corrective Forcing (CoF), a post‑training method that aligns diffusion and flow generative models for speech enhancement by training them on self‑generated rollout states. CoF corrects predictions toward ground truth under dynamic sampling schedules and regularizes local evolution with counterfactual transitions, applying a unified objective across both model types. Experiments on SB‑VE and OT‑CFM show improved perceptual quality, reconstruction fidelity, and robustness to varying sampling steps.

By Qing Yao, Lijian Gao, Qirong Mao
arXiv AI
Sep 4

Test-time adaptation for speech enhancement with an autoregressive speech prior

The paper proposes a single‑utterance test‑time adaptation (TTA) method for speech enhancement that uses an autoregressive prior trained on clean speech latent representations from a neural audio codec. The adaptation regularizes a pretrained enhancement model by minimizing the Kullback‑Leibler divergence between the enhanced speech distribution and the clean speech prior. Experiments on multiple noisy speech datasets demonstrate consistent improvements in speech quality, especially when training and testing noise conditions differ.

By Sofiene Kammoun, Simon Leglaive, Xavier Alameda-Pineda, Timo Gerkmann
arXiv Machine Learning
Sep 15

The Limits of Reference-Free Speech Quality Metrics as Evaluators and Rewards on Modern Text-to-Speech

arXiv:2609.13150v1 Announce Type: cross Abstract: Reference-free quality predictors such as UTMOS, DNSMOS and SCOREQ are the de facto automatic evaluators for text-to-speech (TTS) and are increasingl...

By Antonis Asonitis, Juan Pablo Zuluaga Gomez, Francesco Verdini, Aref Farhadipour, Marzieh Razavi, Pierre-Edouard Honnet, Vijeta Avijeet
arXiv AI
Jun 9

End-to-End Training for Discrete Token LLM based TTS System

arXiv:2606. 09234v1 Announce Type: cross Abstract: Recent state-of-the-art (SOTA) text-to-speech (TTS) systems typically adopt a cascaded pipeline consisting of a speech tokenizer, an autoregressive large language model (LLM), and a diffusion based flow-matching (FM) model, with these components trained independently.

By Changfeng Gao, Yong Ren, Jun Yuan, Ye Bai, Zhao You, ShiDong Shang