arXiv:2604. 08558v2 Announce Type: replace-cross Abstract: Recent decoder-only autoregressive text-to-speech (AR-TTS) models produce high-fidelity speech, but their memory and compute costs scale quadratically with sequence length due to full self-attention.
By Hanna Lee, Tan Dat Nguyen, Jaehoon Kang, Kyuhong Shim
arXiv:2601. 19919v2 Announce Type: replace-cross Abstract: Knowledge distillation (KD) is one of the most effective paradigms for compressing large-scale foundation models into deployable architectures.
By Junseok Lee, Nahun Kim, Sangyong Lee, Chang-Jae Chun
The paper introduces acoustic-to-text KV compression for full‑duplex speech models, converting acoustic key‑value states into compact textual memory during listening‑time slack. When the KV cache exceeds a target budget, older acoustic states are evicted while transcripts and recent acoustic context are retained. Experiments on ten‑minute LongSpeech sessions show a 64.6% reduction in peak streaming KV‑cache size and improved transcription, temporal question answering, and summarization, with comparable pause‑handling, turn‑taking, and interruption performance in Full‑Duplex‑Bench.
By Yejin Lee, Seungbeom Kim, Yongha Lee, Kyuhong Shim
arXiv:2606. 11766v1 Announce Type: cross Abstract: Distilling a large speech foundation model (SFM) into an efficient student model has been successfully applied to low-resource environments.
By Eungbeom Kim, Kyogu Lee
REALM is a retrospective knowledge distillation framework that enables causal decoding of behavior from local field potentials (LFPs). It trains a bidirectional Mamba‑2 teacher on multi‑session data using continuous masked autoencoding, then distills its representations into a compact causal student model. The resulting LFP‑only decoder achieves the highest mean accuracy among compared methods, surpassing state‑of‑the‑art baselines while using fewer parameters and less pretraining time.
By Peicheng Wu, Zhenyu Bu, Runze Ma, Lin Du
arXiv:2608. 08638v1 Announce Type: cross Abstract: Zero-shot text-to-speech (TTS) now supports interactive assistants, personalized media, and accessibility tools.
By Yuqian Zhang, Yao Shi, Kexin Huang, Botian Jiang, Zhe Xu, Yiwei Zhao, Min Liang, Shuang Chen, Xipeng Qiu