arXiv:2511. 20973v2 Announce Type: replace-cross Abstract: Large Audio Language Models (LALMs) deliver strong performance across speech and audio tasks, but their audio encoders generate high-rate token sequences (e.
By Saurabhchand Bhati, Samuel Thomas, Hilde Kuehne, Rogerio Feris, James Glass
arXiv:2607. 29363v1 Announce Type: cross Abstract: Balancing sequence length, representational capacity, and long-horizon stability is a central problem in autoregressive (AR) speech and audio generation.
By Yi Luo, Rongzhi Gu, Jixun Yao
arXiv:2606. 09019v1 Announce Type: cross Abstract: Codec-based autoregressive (AR) speech language models have achieved strong text-to-speech (TTS) quality by modeling speech as sequences of discrete audio tokens with large pretrained backbones.
By Yejin Lee, Junwon Moon, Hyoeun Kim, Hyunjin Choi, Heeseung Kim, Kyuhong Shim
arXiv:2412. 11449v2 Announce Type: replace-cross Abstract: We propose WHISPER-GPT: A generative large language model (LLM) for speech and music that allows us to work with continuous audio representations and discrete tokens simultaneously as part of a single architecture.
By Prateek Verma
arXiv:2609.22851v2 Announce Type: cross
Abstract: Large Audio Language Models (LALMs) utilize either continuous features or discrete tokens, yet the optimal representation paradigm for general audio...
By Jing Peng, Zichao Nie, Zhisheng Zhang, Jingran Xie, Zhiyong Wu
arXiv:2607. 13013v1 Announce Type: new Abstract: Automatic speech recognition is dominated by autoregressive decoders that emit one token at a time.
By Harsha Vardhan Khurdula, Abhinav Kumar Singh, Yoeven D Khemlani, Vineet Agarwal
Automatic speech recognition is dominated by autoregressive decoders that emit one token at a time. We ask whether a discrete diffusion language model can transcribe speech instead, refining a whole transcript in parallel over a small number of denoising steps.
The paper introduces geometric iterative retrieval, a new approach for resynthesizing high‑quality audio from coarse Residual Vector Quantization (RVQ) codec tokens. Instead of choosing between discrete token prediction or continuous regression, the method performs contrastive retrieval within the continuous codebook space, leveraging the RVQ hierarchy as an iterative decomposition. Experiments on speech and music codec restoration tasks demonstrate that this technique outperforms both single‑pass token prediction and one‑step regression baselines.
By Leo Schmidt-Traub, Fr\'ed\'eric Berdoz, Luca A. Lanzend\"orfer, Roger Wattenhofer
TokenMapper is a framework that enables direct translation between different speech tokenizers, allowing heterogeneous speech models to communicate without converting tokens to waveform audio. It handles mismatched token spaces, including single and multi-codebook representations, while maintaining a shared effective token rate. Experiments on GLM-4-Voice, MiMi, and DualCodec show that TokenMapper achieves word error rates close to native reconstructions, comparable human MOS scores, and significantly reduces latency compared to waveform bridging.
By Tal Kozakov, Tal Rosenwein, Eliya Nachmani
arXiv:2606. 02739v1 Announce Type: cross Abstract: Audio tokenizers serve as the discrete interface between continuous audio and Audio Language Models (ALMs), but existing tokenizers often struggle to support both understanding and generation.
By Hui Li, Yangfan Gao, Junlin Shang, Changhao Jiang, Tao Gui, Qi Zhang, Xuanjing Huang
arXiv:2609.13909v1 Announce Type: cross
Abstract: Continuous-latent Autoregressive Diffusion Transformer (AR-DiT) models have demonstrated immense potential in zero-shot speech generation. However, t...
By Ziyu Zhang, Tianlun Zuo, Hanzhao Li, Haoyu Zhang, Lei Xie
arXiv:2606. 10233v1 Announce Type: cross Abstract: While speech quality is typically assessed on complete utterances, streaming and generative systems require incremental estimation from partial audio.
By Zhuoyan Tao, Jiatong Shi, Hye-jin Shim, Shinji Watanabe