The paper adapts Reinforce Adjoint Matching (RAM) to generative speech enhancement, allowing a pretrained model to be post‑trained on real recordings using weak supervision such as text transcripts. RAM shifts the model’s conditional distribution toward higher‑reward outputs by generating enhanced speech on‑policy, evaluating each output with a potentially non‑differentiable reward, and analytically re‑noising the endpoint to create inputs for a reward‑guided regression objective. Experiments on real CHiME‑4 recordings show a 5.08‑percentage‑point reduction in word error rate compared to the pretrained FlowSE model, while maintaining all reported non‑intrusive speech quality metrics and receiving no significant preference in a subjective listening test.
By Julius Richter, Christoph Boeddeker, Yoshiki Masuyama, Kohei Saijo, Dominik Klement, Gordon Wichern, Jonathan Le Roux
arXiv:2609.13045v1 Announce Type: new
Abstract: Speech-to-speech translation (S2ST) has advanced significantly with speech LLMs, offering the potential for joint optimization and preserving non-lingu...
By Hayato Futami, Hassan Shahmohammadi, Tushar Dhyani, Alkis Koudounas, Rapha\"el Lafargue, Yosuke Kashiwagi, Quentin Jodelet, Emiru Tsunoo
arXiv:2512. 20978v2 Announce Type: replace-cross Abstract: Language Model (LM)-based generative modeling has emerged as a promising direction for TSE, offering potential for improved generalization and high-fidelity speech.
By Haoyang Li, Xuyi Zhuang, Azmat Adnan, Ye Ni, Wei Rao, Shreyas Gopal, Eng Siong Chng, Boon Siew Han, Yuanjin Zheng
The paper proposes a single‑utterance test‑time adaptation (TTA) method for speech enhancement that uses an autoregressive prior trained on clean speech latent representations from a neural audio codec. The adaptation regularizes a pretrained enhancement model by minimizing the Kullback‑Leibler divergence between the enhanced speech distribution and the clean speech prior. Experiments on multiple noisy speech datasets demonstrate consistent improvements in speech quality, especially when training and testing noise conditions differ.
By Sofiene Kammoun, Simon Leglaive, Xavier Alameda-Pineda, Timo Gerkmann
arXiv:2604. 14606v2 Announce Type: cross Abstract: Universal speech enhancement (USE) aims to restore speech signals from diverse distortions across multiple sampling rates.
By Xiaobin Rong, Zheng Wang, Yushi Wang, Jun Gao, Jing Lu
arXiv:2603. 13952v3 Announce Type: replace-cross Abstract: In existing Audio-Visual Speech Enhancement (AVSE) methods, objectives such as Scale-Invariant Signal-to-Noise Ratio (SI-SNR) and Mean Squared Error (MSE) are widely used; however, their correlation with perceived speech quality is often suboptimal and provides limited interpretability for optimization.
By Chih-Ning Chen, Jen-Cheng Hou, Hsin-Min Wang, Shao-Yi Chien, Yu Tsao, Fan-Gang Zeng
arXiv:2609.36577v1 Announce Type: cross
Abstract: Audio large language models (ALLMs) can reason about the content of audio recordings to perform complex tasks. However, these capabilities usually co...
By Zhenhong Zhou, Xuanyue Zhao, Youji Liu, Yuanhe Zhang, Xiaoyu Ma, Lianyu Hu, Yang Liu
arXiv:2608. 09288v1 Announce Type: cross Abstract: Audio-visual speech enhancement under real-world conditions remains challenging due to unreliable visual inputs and the lack of large-scale training data with realistic acoustic conditions.
By Wei Zhou, Wanyi Ning, Yinshang Guo, Qianxiao Fang, Haitao Qian, Yingpeng Li
arXiv:2604. 01832v1 Announce Type: cross Abstract: We introduce GAP-URGENet, a generative-predictive fusion framework developed for Track 1 of the ICASSP 2026 URGENT Challenge.
By Xiaobin Rong, Yushi Wang, Zheng Wang, Jing Lu
arXiv:2607. 10191v1 Announce Type: cross Abstract: Generative streaming models for Target Speaker Extraction (TSE) commonly exhibit a quality--intelligibility trade-off, wherein naive optimization for perceptual audio quality tends to degrade speech intelligibility, and conversely.
By Shuhai Peng, Jinjiang Liu, Hui Lu, Liyang Chen, Guiping Zhong, Jiakui Li, Shiyin Kang, Zhiyong Wu
arXiv:2601.09050v2 Announce Type: replace
Abstract: Tonal low-resource languages are widely spoken but remain underserved by modern speech technologies. A central challenge is learning speech represe...
By Tianyi Xu, Xuan Ouyang, Binwei Yao, Shoua Xiong, Sara Misurelli, Maichou Lor, Junjie Hu
arXiv:2607. 08111v1 Announce Type: cross Abstract: Training target speaker extraction (TSE) models for real conversational mixtures remains challenging because large-scale training corpora and clean target speech for supervision are unavailable.
By Wanyi Ning, Wei Zhou, Yingpeng Li, Yinshang Guo, Haitao Qian, Yiming Cheng