arXiv:2603. 15590v2 Announce Type: replace Abstract: There have been numerous attempts to distill quadratic attention-based large language models (LLMs) into sub-quadratic linearized architectures.
By Lukas Hauzenberger, Niklas Schmidinger, Thomas Schmied, Anamaria-Roberta Hartl, David Stap, Pieter-Jan Hoedt, Maximilian Beck, Sebastian B\"ock, G\"unter Klambauer, Sepp Hochreiter
arXiv:2609.07086v1 Announce Type: new
Abstract: Transformers are a dominant architecture in modern machine learning, powering applications across vision, language, and beyond. At the core of their su...
By Hemanth Saratchandran, Simon Lucey
arXiv:2603. 26556v2 Announce Type: replace-cross Abstract: Converting a pretrained Transformer into a more efficient hybrid model through distillation offers a promising approach to reducing inference costs.
By Juan Gabriel Kostelec, Qinghai Guo
arXiv:2601. 11667v2 Announce Type: replace-cross Abstract: Transformer architectures deliver state-of-the-art accuracy via dense full-attention, but their quadratic time and memory complexity with respect to sequence length limits practical deployment.
By Xiaojie Xia, Huigang Zhang, Chaoliang Zhong, Jun Sun, Yusuke Oishi
arXiv:2609.38342v1 Announce Type: new
Abstract: On-policy self-distillation uses a model as its own teacher to provide dense supervision for reasoning, often through reference-solution conditioning....
By Zhexi Lu, Subhajit Chaudhury, Tejaswini Pedapati, Keerthiram Murugesan, Lei Yu
arXiv:2607. 19358v1 Announce Type: new Abstract: Recent advances in long chain-of-thought reasoning models such as DeepSeek-R1 have led to increasingly longer inference context lengths under the test-time scaling paradigm.
By Yu Zhao, Zekun Zhang, Fan Jiang, Bo Zeng, Linlong Xu, Shimin Shan, Yu Liu, Longyue Wang, Weihua Luo
arXiv:2510.19266v3 Announce Type: replace
Abstract: State-space models (SSMs) have emerged as promising alternatives to Transformers for sequence modeling. However, training competitive SSMs from scr...
By Penghao Wang, Yuhao Zhou, Mengxuan Wu, Panpan Zhang, Zhangyang Wang, Kai Wang
arXiv:2604. 08558v2 Announce Type: replace-cross Abstract: Recent decoder-only autoregressive text-to-speech (AR-TTS) models produce high-fidelity speech, but their memory and compute costs scale quadratically with sequence length due to full self-attention.
By Hanna Lee, Tan Dat Nguyen, Jaehoon Kang, Kyuhong Shim
The paper introduces On-Policy Attention Self-Distillation (OPASD), a method that augments token-level supervision with solution-conditioned attention distillation for reasoning models. OPASD projects a privileged teacher’s attention onto student-visible positions, renormalizes the distribution, and aligns it with the student. Experiments on three model sizes and four math benchmarks show that OPASD improves accuracy by 4.98–8.40 percentage points, reduces generated tokens by 73.9%, cuts compute by 72.6%, and trains 1.53× faster compared to token-only distillation.
By Safaeid Hossain Arib, Rabeya Akter, Ismam Nur Swapnil, Md. Faiyaz Abdullah Sayeedi, Tasnim Mohiuddin, Md Mofijul Islam
Video DeltaNet (VDN) introduces a hybrid attention mechanism for video diffusion models, combining local Softmax attention with a bidirectional linear memory branch called Video Delta Attention (VDA). The design updates memory once per frame, uses separate output projections and learnable gates to balance the two branches, and employs a staged teacher‑alignment recipe to integrate the new pathway into pretrained models. When applied to MiniMax H3, VDN achieves a 14.5× speedup, completing 14.3‑second, 768p video denoising in 6.70 seconds on eight NVIDIA B200 GPUs compared to the 50‑step dense baseline.
The paper introduces AnchorCache, a parameter‑free token‑layout and attention‑mask design that decouples reference tokens from the target in in‑context diffusion transformers. By inserting static text anchors, the method conditions reference representations on the instruction during cache construction, enabling exact key‑value reuse across denoising steps. To restore quality lost by this structural change, the authors employ teacher‑forced velocity distillation followed by a brief on‑policy stage, achieving full‑attention quality while delivering up to 6.40× speedup in diffusion transformer inference across image, speech, and video benchmarks.
By Yangshuai Liu, Zheming Li, Jiaao Li, Kang He, Ziliang Lai, Zhitai Liu, Chengru Song
arXiv:2608. 01263v1 Announce Type: new Abstract: On-policy distillation (OPD) samples trajectories from the current student policy and minimizes token-level divergence between student and teacher next-token distributions at prefixes along those trajectories.
By Leyan Xue, Feng Xiong, Mingjun Ma, Changqing Zhang