arXiv Machine Learning

Geometric Iterative Retrieval for Neural Audio Codec Resynthesis

The paper introduces geometric iterative retrieval, a new approach for resynthesizing high‑quality audio from coarse Residual Vector Quantization (RVQ) codec tokens. Instead of choosing between discrete token prediction or continuous regression, the method performs contrastive retrieval within the continuous codebook space, leveraging the RVQ hierarchy as an iterative decomposition. Experiments on speech and music codec restoration tasks demonstrate that this technique outperforms both single‑pass token prediction and one‑step regression baselines.

Hugging Face Trending Papers
Jun 3

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding

Neural audio codecs are a key component of speech processing pipelines, compressing audio into discrete tokens for downstream modeling. However, existing codecs struggle to balance reconstruction quality with token efficiency, often encoding perceptually irrelevant information such as background noise and recording artifacts at the expense of linguistically and acoustically meaningful content.

arXiv Machine Learning
Aug 27

BRIDLE: Generalized Self-supervised Learning with Quantization

BRIDLE is a self‑supervised encoder pretraining framework that extends bidirectional training to audio, image, and video by incorporating residual quantization (RQ) with multiple hierarchical codebooks. This approach allows fine‑grained discretization of latent representations and interleaves training between the encoder and tokenizer. Experiments show that BRIDLE achieves state‑of‑the‑art results on audio classification benchmarks and competitive performance on image and video classification tasks, outperforming traditional vector‑quantization methods.

By Hoang M. Nguyen, Satya N. Shukla, Qiang Zhang, Hanchao Yu, Sreya D. Roy, Dipesh Tamboli, Taipeng Tian, Lingjiong Zhu, Yuchen Liu