arXiv:2608.22898v1 Announce Type: new
Abstract: Diffusion language models (DLMs) alleviate the inherent latency bottleneck of autoregressive (AR) large language models (LLMs), but their degraded gene...
By Hyeongsoo Lim, Jinyoung Kim, Eunseo Seo, Minho Jang, Jiwon Yoon
arXiv:2607.22629v3 Announce Type: replace
Abstract: Large Reasoning Models produce long, explicit chains of intermediate steps before generating a final answer at inference time. These intermediate t...
By Durgesh Kalwar, Vardhan Palod, Jaya Adithya Pavuluri, Subbarao Kambhampati
arXiv:2605. 15532v3 Announce Type: replace-cross Abstract: Distillation enables compact Vision-Language Models (VLMs) to obtain strong reasoning capabilities, yet the prompts driving this process are typically chosen via simple heuristics or aggregated from off-the-shelf datasets.
By Jaehun Jung, Hyunwoo Kim, Brandon Cui, Ximing Lu, David Acuna, Prithviraj Ammanabrolu, Yejin Choi
arXiv:2402. 14035v4 Announce Type: replace-cross Abstract: Knowledge distillation from foundation models to compact domain models is challenging due to substantial gaps in capacity, architecture, and modality.
By Zichang Liu, Qingyun Liu, Yuening Li, Liang Liu, Anshumali Shrivastava, Shuchao Bi, Lichan Hong, Ed H. Chi, Zhe Zhao
arXiv:2609.36734v1 Announce Type: new
Abstract: Knowledge Distillation (KD) trains a smaller-capacity student model to imitate a larger-capacity teacher model by matching output distributions, implic...
By Ayan Sengupta, Vaibhav Seth, Tanmoy Chakraborty
arXiv:2604. 08558v2 Announce Type: replace-cross Abstract: Recent decoder-only autoregressive text-to-speech (AR-TTS) models produce high-fidelity speech, but their memory and compute costs scale quadratically with sequence length due to full self-attention.
By Hanna Lee, Tan Dat Nguyen, Jaehoon Kang, Kyuhong Shim
arXiv:2608. 01672v1 Announce Type: cross Abstract: Effective long-context modeling is not merely about retaining more of the past, but about preserving the information that may prove relevant later.
By Zixuan Wang, Xingyu Dang, Rui-Jie Zhu, Zixin Wen, Hengyu Fu, Wenhao Chai, Jason D. Lee
arXiv:2407. 13911v5 Announce Type: replace-cross Abstract: Prompt-based continual learning has shown strong performance in rehearsal-free class-incremental learning by adapting learnable prompts while freezing a pre-trained Vision Transformer (ViT) backbone.
By Qifan Zhang, Yunhui Guo, Yu Xiang
RT-SEMamba is a fully causal speech enhancement model that uses causal time‑frequency Mamba blocks instead of Transformer‑based architectures, allowing efficient long‑form inference with a fixed‑size recurrent state. The authors introduce a progressive knowledge distillation strategy that compresses an 8‑layer teacher into a single‑layer student by jointly distilling spectral outputs and intermediate representations. On the Voicebank‑DEMAND benchmark, the 8‑layer model achieves 3.32 PESQ under a 25 ms latency constraint, while the distilled 1‑layer student improves from 3.06 to 3.18 PESQ, maintains the same steady‑state real‑time factor, and runs 2.64× faster than the teacher.
By Rong Chao, Sung-Feng Huang, Moreno La Quatra, Sabato Marco Siniscalchi, Wen-Huang Cheng, Szu-Wei Fu, Yu Tsao
arXiv:2606. 06712v1 Announce Type: cross Abstract: We study the transformation of autoregressive models (ARLMs) into diffusion language models (DLMs).
By Xingyu Su, Jacob Helwig, Shubham Parashar, Atharv Chagi, Lakshmi Jotsna, Degui Zhi, James Caverlee, Dileep Kalathil, Shuiwang Ji
arXiv:2607. 02502v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) has emerged as a practical method for training large language models (LLMs) to reason, where a single model acts as both the teacher and the student with different levels of information access.
By Yunhe Li, Hao Shi, Wenhao Liu, Mengzhe Ruan, Hanxu Hou, Zhongxiang Dai, Shuang Qiu, Linqi Song
The paper introduces DKL, a method that decouples knowledge learning from instruction tuning in language models. Instead of fine‑tuning the instruction‑tuned model directly, DKL first extends pre‑training on a base model to embed new knowledge, then merges these weights into the instruction‑tuned model, preserving its instruction‑following abilities. Experiments show DKL raises RAG accuracy from 54.17 % to 79.26 % on retrieval failures, outperforming prior methods while using far less training data.