arXiv:2607. 14952v1 Announce Type: new Abstract: A growing gap separates inference context lengths from RL post-training: inference systems are approaching million-token contexts, while post-training workloads often remain at 256K tokens or below and rely on length generalization at deployment.
By Changhai Zhou, Kieran Liu, Yuhua Zhou, Qian Qiao, Jun Gao, Harry Zhang, Irvine Lu, Nolan Ho, Lucian Li, Andrew Lei, Cleon Cheng, Steven Chiang, Yihang Zeng, Di Zhang, Rio Yang, Kaijie Chen, Andrew Chen, Pony Ma, Weizhong Zhang, Cheng Jin
arXiv:2609.24197v1 Announce Type: new
Abstract: Speculative decoding losslessly accelerates large language model inference by having a lightweight draft model predict future tokens for verification b...
By Weifan Jiang, Krishna Teja Chitty-Venkata, Megan Flynn, Reed Meyerson, Zhenting Qi, Tianyu Wu, Eldar Kurtic, Minlan Yu, Alexandre Marques
The paper introduces On‑Demand Attention (ODA), a decoding strategy that lets pretrained language models decide when to use global attention based on a lightweight recall head. ODA keeps the original model weights unchanged, only training the recall head, and can be implemented with GPU‑side conditional execution to reduce global reads. Experiments on Qwen, Gemma, and hybrid‑attention models show that ODA largely recovers performance lost by local attention while cutting the number of global attention operations.
By Haibo Feng, Ruiqi Liang, Hanyang Peng, Shiqi Yu
arXiv:2606. 00144v1 Announce Type: cross Abstract: Speculative decoding speeds up autoregressive decoding by using a drafter to propose multiple tokens that a verifier validates in parallel.
By Liang He, Jingbo Wen, Qishi Zhan, Yixiong Chen, Kangning Cui, Qizhen Lan, Xilu Wang
arXiv:2606. 18967v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a representative post-training paradigm for LLMs, enabling strong reasoning and agentic capabilities.
By Minseo Kim, Minjae Lee, Seunghyuk Oh, Kevin Galim, Donghoon Kim, Coleman Hooper, Harman Singh, Amir Gholami, Hyung Il Koo, Wonjun Kang
arXiv:2608. 04962v1 Announce Type: new Abstract: Reinforcement learning (RL) post-training improves the reasoning capabilities of large language models, but autoregressive rollout generation remains a major efficiency bottleneck.
By Nhat Minh Pham, Duy Tung Doan, Thi Duyen Ngo, Vinh Van Nguyen, Khac-Hoai Nam Bui
GrowMTP is a method that trains a draft head entirely within the reinforcement learning (RL) loop, using supervision from the RL verification step and a rollout distribution that is narrower than pretraining. By detaching draft‑head updates from the policy backbone, it enables online training of the draft head from scratch. Experiments on Qwen3‑4B, MiMo‑7B‑SFT, and Qwen3.5‑4B‑Base show rollout speedups ranging from 1.36× to 2.13× and overall end‑to‑end speedups from 1.20× to 1.60×, making it a modular acceleration component for RL frameworks lacking pretrained draft heads.
By Minghua He, Lingzhe Zhang, Yuan Liu, Xiao Zhou, Aiwei Liu
TIDE (Temporal Incremental Draft Engine) is a serving‑engine‑native framework that integrates online draft adaptation into high‑performance LLM inference. By reusing intermediate hidden states from the target model as training signals, TIDE avoids extra target model computation and serving‑time overhead, activating speculation and draft training only when beneficial. On heterogeneous GPU clusters, TIDE achieves up to 1.66× higher throughput than no‑speculation baselines, reduces training time by up to 3.02×, cuts storage needs by 24×, and improves system throughput by up to 1.22×.
By Jiyoung Park, Hankyu Jang, Changseok Song, Wookeun Jung
arXiv:2609.24698v1 Announce Type: new
Abstract: Repeated execution of the target model during autoregressive decoding is a major source of LLM inference latency. Unlike linear speculation, which foll...
By Changxu Liu, Zhaogeng Li
arXiv:2608. 01651v1 Announce Type: cross Abstract: Hybrid-attention large language models combine full attention with recurrent linear attention to reduce long-context inference costs, yet their autoregressive decoding remains memory-bound.
By Li Wang, Yi Su, Xiabao Wu, Chiran You, Yongchao Liu, Zhan Qiu, Juelu Zhang, Jiajun Zheng, Fangxin Liu, Jie Zhang, Chen Tian, Chengying Huan
The paper introduces On‑Demand Attention (ODA), a local‑first decoding strategy that predicts when a pretrained language model would benefit from global attention. By training only a lightweight recall head, ODA selectively triggers global attention during generation, keeping pretrained weights unchanged and preserving the full key‑value cache for future recall. Experiments on Qwen, Gemma, and hybrid‑attention models show that ODA largely recovers performance lost with local attention while significantly cutting global reads, enabling faster long‑context inference.
arXiv:2607. 21535v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive generation by having a cheap draft propose tokens that a target verifies in parallel.
By Alagappan Valliappan