The paper adapts Reinforce Adjoint Matching (RAM) to generative speech enhancement, allowing a pretrained model to be post‑trained on real recordings using weak supervision such as text transcripts. RAM shifts the model’s conditional distribution toward higher‑reward outputs by generating enhanced speech on‑policy, evaluating each output with a potentially non‑differentiable reward, and analytically re‑noising the endpoint to create inputs for a reward‑guided regression objective. Experiments on real CHiME‑4 recordings show a 5.08‑percentage‑point reduction in word error rate compared to the pretrained FlowSE model, while maintaining all reported non‑intrusive speech quality metrics and receiving no significant preference in a subjective listening test.
By Julius Richter, Christoph Boeddeker, Yoshiki Masuyama, Kohei Saijo, Dominik Klement, Gordon Wichern, Jonathan Le Roux
arXiv:2510. 20441v2 Announce Type: replace-cross Abstract: Neural audio codecs have largely promoted the application of language models (LMs) for speech applications.
By Haoyin Yan, Chengwei Liu, Shaofei Xue, Xiaotao Liang, Yinghao Liu, Yuxiang Kong, Zheng Xue
StrixAE is an intelligent audio enhancement agent that uses a multimodal large language model to coordinate multiple enhancement and personalization models. It is trained in two stages: first, chain‑of‑thought supervised fine‑tuning on AcoustBench, and second, Audio Perception Reinforcement Learning that optimizes format validity, structural coherence, and perceptual quality. The approach yields state‑of‑the‑art performance and strong generalization on real‑world test datasets, outperforming many existing open‑source and proprietary solutions.
By Chenglin Wu, Junjie Wu, Jinhang Chen, Mingyang Chen, Zixu Lin, Jiabian Chen, Xinghao Ding, Xiaotong Tu
arXiv:2608. 09930v1 Announce Type: cross Abstract: Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive.
By Oluwanifemi Bamgbose, Simon Rosen, Jash Shah, Lindsay Devon Brin, Hoang H Nguyen, Anke Koelzer, Rachel Hansen, Tara Bogavelli, Fanny Riols
arXiv:2608. 18607v2 Announce Type: replace Abstract: Using reinforcement learning to post-train joint video-audio generation models requires a reward signal.
By Yinming Huang, Shuyuan Tu, Xi Yan, Zihan Yang, Jianhua Han, Xu Hang, Yu-Gang Jiang, Zuxuan Wu
arXiv:2509. 14659v3 Announce Type: replace-cross Abstract: Current audio captioning relies on supervised learning with paired audio-caption data, which is costly to curate and may not reflect human preferences in real-world scenarios.
By Kartik Hegde, Rehana Mahfuz, Yinyi Guo, Erik Visser
arXiv:2606.28249v2 Announce Type: replace-cross
Abstract: Recently, Large Language Model (LLM)-based Text-to-Speech (TTS) models have achieved remarkable naturalness. However, the standard Supervised...
By Sihang Nie, Xiaofen Xing, Rui Xing, Haoming Li, Ruitong Xiao, Jingyuan Xing, Baiji Liu, Xiangmin Xu
arXiv:2608.21176v1 Announce Type: cross
Abstract: Automatic speech quality assessment aims to predict Mean Opinion Scores (MOS) consistent with human subjective perception and is essential for evalua...
By Naiyuan Li, Li Dong, Diqun Yan
The paper proposes a single‑utterance test‑time adaptation (TTA) method for speech enhancement that uses an autoregressive prior trained on clean speech latent representations from a neural audio codec. The adaptation regularizes a pretrained enhancement model by minimizing the Kullback‑Leibler divergence between the enhanced speech distribution and the clean speech prior. Experiments on multiple noisy speech datasets demonstrate consistent improvements in speech quality, especially when training and testing noise conditions differ.
By Sofiene Kammoun, Simon Leglaive, Xavier Alameda-Pineda, Timo Gerkmann
The paper introduces a method to improve audio‑visual speech recognition by applying contrastive decoding (CD) that contrasts audio‑only with audio‑visual conditioning within the same model. It addresses the issue of a fixed CD strength by scaling the influence adaptively for each token, using reliability signals from attention dynamics and predictive divergence. Experiments on the LRS3 dataset demonstrate consistent gains in both clean and low‑SNR scenarios.
By YoungChae Kim, Da-Hee Yang, Joon-Hyuk Chang
arXiv:2608. 09288v1 Announce Type: cross Abstract: Audio-visual speech enhancement under real-world conditions remains challenging due to unreliable visual inputs and the lack of large-scale training data with realistic acoustic conditions.
By Wei Zhou, Wanyi Ning, Yinshang Guo, Qianxiao Fang, Haitao Qian, Yingpeng Li
arXiv:2609.30631v1 Announce Type: cross
Abstract: Online audio-visual target-speaker extraction aims to remove competing voices while preserving speech quality and bounding lookahead. Existing extrac...
By Rayhan Rashed, Senja Filipi, Ross Cutler