arXiv AI By Sofiene Kammoun, Simon Leglaive, Xavier Alameda-Pineda, Timo Gerkmann

Test-time adaptation for speech enhancement with an autoregressive speech prior

Read the original on arXiv AI →

The paper proposes a single‑utterance test‑time adaptation (TTA) method for speech enhancement that uses an autoregressive prior trained on clean speech latent representations from a neural audio codec. The adaptation regularizes a pretrained enhancement model by minimizing the Kullback‑Leibler divergence between the enhanced speech distribution and the clean speech prior. Experiments on multiple noisy speech datasets demonstrate consistent improvements in speech quality, especially when training and testing noise conditions differ.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 9

Speech Enhancement Based on Drifting Models

arXiv:2604. 24199v4 Announce Type: replace-cross Abstract: We propose Speech Enhancement based on Drifting Models (DriftSE), a novel generative framework that formulates denoising as an equilibrium problem.

By Liang Xu, Diego Caviedes-Nozal, W. Bastiaan Kleijn, Longfei Felix Yan, Rasmus Kongsgaard Olsson
arXiv AI
Sep 4

Masked Autoregressive Speech Enhancement with Continuous Neural Audio Codec Representations

The paper introduces Masked Autoregressive Speech Enhancement (MARSE), a method that iteratively decodes masked clean speech frames using continuous latent representations from a neural audio codec (DAC). Unlike prior approaches that relied on discrete token representations, MARSE employs a Conformer model and explores various decoding policies to balance speech enhancement performance with computational cost. The authors provide audio examples and code online to demonstrate the method’s effectiveness.

By Yoto Fujita, Simon Leglaive, Laurent Girin
arXiv AI
Sep 2

Cleaner Speech, Weaker Generalization: Revisiting Pitt-Derived Benchmarks for Alzheimer's Disease Detection

The study examines how speech preprocessing—such as enhancement, sample selection, and demographic balancing—affects Alzheimer’s disease detection models that use the Pitt Corpus. Experiments reveal that while speech‑enhanced datasets boost in‑domain accuracy, they diminish cross‑dataset robustness and introduce class imbalance and prediction shifts, even when training and testing enhancements are matched. Large audio‑language models show similar sensitivity, indicating that cleaner speech does not guarantee better real‑world performance.

By Luqi Sun, Shreeram Suresh Chandra, Lin Zhang, You-Jin Li, Brian MacWhinney, Yu Tsao, Emily Mower Provost, Berrak Sisman