arXiv AI By Chih-Ning Chen, Jen-Cheng Hou, Hsin-Min Wang, Shao-Yi Chien, Yu Tsao, Fan-Gang Zeng

LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement

Read the original on arXiv AI →

arXiv:2603. 13952v3 Announce Type: replace-cross Abstract: In existing Audio-Visual Speech Enhancement (AVSE) methods, objectives such as Scale-Invariant Signal-to-Noise Ratio (SI-SNR) and Mean Squared Error (MSE) are widely used; however, their correlation with perceived speech quality is often suboptimal and provides limited interpretability for optimization.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 25

Transcript-Supervised Post-Training of Generative Speech Enhancement on Real Recordings via Reinforce Adjoint Matching

The paper adapts Reinforce Adjoint Matching (RAM) to generative speech enhancement, allowing a pretrained model to be post‑trained on real recordings using weak supervision such as text transcripts. RAM shifts the model’s conditional distribution toward higher‑reward outputs by generating enhanced speech on‑policy, evaluating each output with a potentially non‑differentiable reward, and analytically re‑noising the endpoint to create inputs for a reward‑guided regression objective. Experiments on real CHiME‑4 recordings show a 5.08‑percentage‑point reduction in word error rate compared to the pretrained FlowSE model, while maintaining all reported non‑intrusive speech quality metrics and receiving no significant preference in a subjective listening test.

By Julius Richter, Christoph Boeddeker, Yoshiki Masuyama, Kohei Saijo, Dominik Klement, Gordon Wichern, Jonathan Le Roux
arXiv AI
Sep 4

StrixAE: An Intelligent Agent for Audio Enhancement under Complex Distortion Coupling in Real-World Scenarios

StrixAE is an intelligent audio enhancement agent that uses a multimodal large language model to coordinate multiple enhancement and personalization models. It is trained in two stages: first, chain‑of‑thought supervised fine‑tuning on AcoustBench, and second, Audio Perception Reinforcement Learning that optimizes format validity, structural coherence, and perceptual quality. The approach yields state‑of‑the‑art performance and strong generalization on real‑world test datasets, outperforming many existing open‑source and proprietary solutions.

By Chenglin Wu, Junjie Wu, Jinhang Chen, Mingyang Chen, Zixu Lin, Jiabian Chen, Xinghao Ding, Xiaotong Tu
arXiv AI
Aug 11

Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions

arXiv:2608. 09930v1 Announce Type: cross Abstract: Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive.

By Oluwanifemi Bamgbose, Simon Rosen, Jash Shah, Lindsay Devon Brin, Hoang H Nguyen, Anke Koelzer, Rachel Hansen, Tara Bogavelli, Fanny Riols