arXiv:2512.21653v2 Announce Type: replace-cross
Abstract: Speech codecs are usually optimized for waveform fidelity, allocating bits to acoustic detail that can be inferred from linguistic structure....
By Liuyang Bai, Weiyi Lu, Li Guo
arXiv:2608. 05727v1 Announce Type: cross Abstract: Neural Audio Codecs are widely adopted in speech generation and editing.
By June Young Yi, Dongwook Lee, Jiheum Yeom, Sungroh Yoon
The paper introduces geometric iterative retrieval, a new approach for resynthesizing high‑quality audio from coarse Residual Vector Quantization (RVQ) codec tokens. Instead of choosing between discrete token prediction or continuous regression, the method performs contrastive retrieval within the continuous codebook space, leveraging the RVQ hierarchy as an iterative decomposition. Experiments on speech and music codec restoration tasks demonstrate that this technique outperforms both single‑pass token prediction and one‑step regression baselines.
By Leo Schmidt-Traub, Fr\'ed\'eric Berdoz, Luca A. Lanzend\"orfer, Roger Wattenhofer
arXiv:2509. 24457v1 Announce Type: cross Abstract: Objective speech-quality metrics are widely used to assess codec performance.
By Wolfgang Mack, Nezih Topaloglu, Laura Lechler, Ivana Bali\'c, Alexandra Craciun, Mansur Yesilbursa, Kamil Wojcicki
arXiv:2602. 15491v2 Announce Type: replace-cross Abstract: Neural audio codecs (NACs) typically encode the short-term energy (gain) and normalized structure (shape) of speech/audio signals jointly within the same latent space.
By Samir Sadok, Laurent Girin, Xavier Alameda-Pineda
arXiv:2608.28916v1 Announce Type: new
Abstract: Automatic speech recognition (ASR) systems are commonly evaluated with word error rate (WER), yet many voice workflows depend on exact written values f...
By Tyler Baumgartner, Brandon Tai, Lisa Kaelin-Martin, Candice Fan, Luc Debaupte, Bill Wang, Yi Zhong
arXiv:2609.36687v1 Announce Type: cross
Abstract: Neural codecs encode continuous signals into compact sequences of discrete tokens, providing an interface for efficient transmission, storage, and to...
By Jihwan Lee, Kleanthis Avramidis, Junhyeok Lee, Tiantian Feng, Najim Dehak, Shrikanth Narayanan
TokenMapper is a framework that enables direct translation between different speech tokenizers, allowing heterogeneous speech models to communicate without converting tokens to waveform audio. It handles mismatched token spaces, including single and multi-codebook representations, while maintaining a shared effective token rate. Experiments on GLM-4-Voice, MiMi, and DualCodec show that TokenMapper achieves word error rates close to native reconstructions, comparable human MOS scores, and significantly reduces latency compared to waveform bridging.
By Tal Kozakov, Tal Rosenwein, Eliya Nachmani
Neural audio codecs are a key component of speech processing pipelines, compressing audio into discrete tokens for downstream modeling. However, existing codecs struggle to balance reconstruction quality with token efficiency, often encoding perceptually irrelevant information such as background noise and recording artifacts at the expense of linguistically and acoustically meaningful content.
The paper introduces Masked Autoregressive Speech Enhancement (MARSE), a method that iteratively decodes masked clean speech frames using continuous latent representations from a neural audio codec (DAC). Unlike prior approaches that relied on discrete token representations, MARSE employs a Conformer model and explores various decoding policies to balance speech enhancement performance with computational cost. The authors provide audio examples and code online to demonstrate the method’s effectiveness.
By Yoto Fujita, Simon Leglaive, Laurent Girin
The study evaluates audio provenance attribution systems, showing that high clean‑benchmark accuracy does not translate to robustness after codec compression. Using a prospectively registered protocol, the authors measured closed‑set attribution performance on two corpora after single‑stage codec transport, finding significant degradation—up to 70.3 Macro‑F1 points for WavLM‑Base+ and 61.0 for W2V2‑BERT 2.0—depending on codec settings and representation. The results demonstrate that clean accuracy alone cannot guarantee deployment robustness across different codecs and representations.
By Gang Shi (Independent Researcher)
arXiv:2606. 02739v1 Announce Type: cross Abstract: Audio tokenizers serve as the discrete interface between continuous audio and Audio Language Models (ALMs), but existing tokenizers often struggle to support both understanding and generation.
By Hui Li, Yangfan Gao, Junlin Shang, Changhao Jiang, Tao Gui, Qi Zhang, Xuanjing Huang