arXiv:2609.14956v1 Announce Type: cross
Abstract: Speech Quality Assessment (SQA) is essential for modern speech technologies, and recent non-intrusive SQA predictors increasingly rely on Speech Foun...
By Alef Iury Siqueira Ferreira, Pedro Lustosa Rege Botelho, Fernanda Silva, Daniel Casanova, Rafael Faustino, Frederico Oliveira, Arlindo Galv\~ao Filho, Anderson da Silva Soares
arXiv:2604. 08558v2 Announce Type: replace-cross Abstract: Recent decoder-only autoregressive text-to-speech (AR-TTS) models produce high-fidelity speech, but their memory and compute costs scale quadratically with sequence length due to full self-attention.
By Hanna Lee, Tan Dat Nguyen, Jaehoon Kang, Kyuhong Shim
arXiv:2601. 19919v2 Announce Type: replace-cross Abstract: Knowledge distillation (KD) is one of the most effective paradigms for compressing large-scale foundation models into deployable architectures.
By Junseok Lee, Nahun Kim, Sangyong Lee, Chang-Jae Chun
RT-SEMamba is a fully causal speech enhancement model that uses causal time‑frequency Mamba blocks instead of Transformer‑based architectures, allowing efficient long‑form inference with a fixed‑size recurrent state. The authors introduce a progressive knowledge distillation strategy that compresses an 8‑layer teacher into a single‑layer student by jointly distilling spectral outputs and intermediate representations. On the Voicebank‑DEMAND benchmark, the 8‑layer model achieves 3.32 PESQ under a 25 ms latency constraint, while the distilled 1‑layer student improves from 3.06 to 3.18 PESQ, maintains the same steady‑state real‑time factor, and runs 2.64× faster than the teacher.
By Rong Chao, Sung-Feng Huang, Moreno La Quatra, Sabato Marco Siniscalchi, Wen-Huang Cheng, Szu-Wei Fu, Yu Tsao
arXiv:2603. 01875v3 Announce Type: replace-cross Abstract: Knowledge distillation (KD) is an essential technique to compress large language models (LLMs) into smaller ones.
By Songming Zhang, Xue Zhang, Tong Zhang, Bojie Hu, Yufeng Chen, Jinan Xu
arXiv:2603. 05121v2 Announce Type: replace-cross Abstract: Speech Large Language Models route speech encoder representations into an LLM decoder that typically accounts for over 90% of total parameters.
By Adel Moumen, Guangzhi Sun, Philip C Woodland