arXiv:2608. 06424v1 Announce Type: cross Abstract: Speech recordings often contain missing, corrupted, or incorrect regions that must be reconstructed or modified without re-synthesizing the entire utterance.
By Iftach Shoham, Tali Dror, Oren Gal, Haim Permuter, Gilad Katz, Eliya Nachmani
arXiv:2606. 11033v1 Announce Type: cross Abstract: Recent efforts to extend large language models (LLMs) to speech inputs typically rely on cascaded ASR-LLM pipelines, end-to-end speech-language models, or bridge/distillation-based adaptation.
By Bo Cheng, Lei Shi, Zhanyu Ma, Yuan Wu, Jun Xu, Jiuchong Gao, Jinghua Hao, Renqing He
arXiv:2608. 02673v1 Announce Type: cross Abstract: Speech editing for content creation requires precise control over both what an edit should do and where it should apply.
By Hankun Wang, Bohan Li, Shi Lian, Xiaoyu Gu, Jing Peng, Da Zheng, Colin Zhang, Kai Yu
arXiv:2606. 20101v3 Announce Type: replace-cross Abstract: Audio editing aims to modify specific content in an existing audio clip according to a text instruction or description while preserving the remaining acoustic content.
By Liting Gao, Yonggang Zhu, Yaru Chen, Dongyu Wang, Shubin Zhang, Zhenbo Li, Jean-Yves Guillemaut, Wenwu Wang
arXiv:2609.03203v3 Announce Type: replace-cross
Abstract: Speech systems increasingly infer how an utterance should be delivered from context, but a plausible delivery plan may not be supported by th...
By Mengzhe Geng
arXiv:2609.13909v1 Announce Type: cross
Abstract: Continuous-latent Autoregressive Diffusion Transformer (AR-DiT) models have demonstrated immense potential in zero-shot speech generation. However, t...
By Ziyu Zhang, Tianlun Zuo, Hanzhao Li, Haoyu Zhang, Lei Xie
AnchorPrompt is an adaptation technique for large audio‑language models that keeps the base model frozen and learns a single block of prompt vectors inserted at the decoder input. By training these prompts through self‑distillation on diverse audio and text perturbations, the method improves answer consistency and reduces hallucinations across multiple benchmarks. The approach is perturbation‑agnostic at inference, enabling zero‑shot transfer to unseen distortions such as reverberation and choice permutations.
By Pooneh Mousavi, Amir Ivry, Mirco Ravanelli, Cem Subakan
arXiv:2606. 31128v1 Announce Type: cross Abstract: Speech editing aims to modify specific portions of an utterance while preserving the remaining speech.
By Chuanbo Zhu, Wuyou Zhou, Rongxiu Zhong, Shilei Zhang, Kun Qian, Yike Guo, Wei Xue
arXiv:2509.10452v3 Announce Type: replace-cross
Abstract: Pretrained automatic speech recognition (ASR) models such as Whisper perform well but still need domain adaptation to handle unseen parlance....
By Akshat Pandey, Karun Kumar, Raphael Tang
arXiv:2609.14344v1 Announce Type: cross
Abstract: Instruction-guided music editors typically process each request independently, limiting their ability to support workflows in which users progressive...
By Quoc-Huy Trinh, Minh-Van Nguyen, Debesh Jha
FireRedAudio is a 9‑billion‑parameter audio language model that separates continuous input representations for audio understanding and speech generation, enabling a single autoregressive LLM to perform tasks such as ASR, zero‑shot TTS, Instruct TTS, and semantic/acoustic speech editing. The model uses a dedicated Audio Encoder for recognition and a RedAE‑based pathway for generation, with the LLM directly generating text or conditioning a flow‑matching DiT to produce acoustic latents. Evaluations show competitive or leading performance in multilingual ASR, content‑accurate zero‑shot TTS, strong instruction following, and significant improvements in speech editing over prior work.
By Feiyu Shen, Fenglong Xie, Junjie Li, Kun Xie, Lei Xie, Xu Tang, Xuelong Geng, Yan Jia, Yao Hu, Yichen Han, Yichen Wu, Ziqi Dai, Junjie Chen, Kai Huang, Manzhen Wei, Yixuan Li
arXiv:2606. 15186v1 Announce Type: cross Abstract: Text-to-audio (TTA) generation has made significant strides, yet achieving precise and consistent audio editing remains a major challenge.
By Yuxuan Jiang, Mingyang Han, Yusheng Dai, Andong Wang, Tianhong Zhou, Jiaxin Ye, Dongxiao Wang, Haoxiang Shi, Boyu Li, Jun Song, Cheng Yu, Bo Zheng, Weibei Dou, Zehua Chen, Jun Zhu