arXiv AI By Samir Sadok, Laurent Girin, Xavier Alameda-Pineda

The Equalizer: Introducing Shape-Gain Decomposition in Neural Audio Codecs

Read the original on arXiv AI →

arXiv:2602. 15491v2 Announce Type: replace-cross Abstract: Neural audio codecs (NACs) typically encode the short-term energy (gain) and normalized structure (shape) of speech/audio signals jointly within the same latent space.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 4

Masked Autoregressive Speech Enhancement with Continuous Neural Audio Codec Representations

The paper introduces Masked Autoregressive Speech Enhancement (MARSE), a method that iteratively decodes masked clean speech frames using continuous latent representations from a neural audio codec (DAC). Unlike prior approaches that relied on discrete token representations, MARSE employs a Conformer model and explores various decoding policies to balance speech enhancement performance with computational cost. The authors provide audio examples and code online to demonstrate the method’s effectiveness.

By Yoto Fujita, Simon Leglaive, Laurent Girin