The paper introduces EpiKV, an epiphany‑aware key–value cache eviction strategy that avoids using the attention matrix. It leverages hidden‑state shifts and recent query–key relevance to rank cached tokens, matching or surpassing the performance of existing attention‑based eviction methods while remaining compatible with fast inference kernels. Experiments on multiple benchmarks show that EpiKV improves inference throughput without sacrificing accuracy.
By Steven Kolawole, Virginia Smith
SIPO (Self‑Instructing Policy Optimization) unifies reinforcement learning with on‑policy self‑distillation by using a contrastive self‑teacher to generate token‑level credit signals. The method samples multiple rollouts per prompt, pairs each with a reference answer and its mistakes, and uses the difference in teacher log‑probabilities to provide dense feedback while still respecting the overall task reward. Experiments on reasoning and code‑generation benchmarks show that SIPO outperforms both RLVR and OPSD baselines without requiring an external teacher or extra generation steps.
By Zhenrui Yue, Huimin Zeng, Yueqi Wang, Yaokun Liu, Fengran Mo, Jinghan Zhang, Mung Yao Jia, Gyuseok Lee, Yang Zhang, Na Wei, Dong Wang
IronLLM-0.6B is a 654‑million‑parameter language model engineered for efficient on‑device inference, featuring a hybrid attention architecture, X‑MTP multi‑token prediction, and a lightweight verification head that yields a 1.48× decoding speedup. Trained on roughly 6.2 trillion tokens with a quality‑oriented pipeline and further refined via Multi‑Domain On‑Policy Distillation, the model adopts an Instruct‑Only design to meet low‑latency requirements. A lighter variant, IronLLM‑0.6B‑Light, replaces RMSNorm with Dynamic Tanh and streamlines costly components to enhance inference and quantization efficiency, offering a strong performance‑efficiency trade‑off for resource‑constrained deployment.
By Changdi Yang, Fengquan Jiao, Haochih Lin, Haoran Yang, Jing Xiao, Liangyu Huo, Suxin Lu, Tiance Chen, Wei Liu, Yinggan Xu, Yunxiang Lu, Zai Zheng, Zhirui Xie, Zhongyang Che, Ziyan Tang, Zuoxiang Zhao, Jian Yao
arXiv:2609.37925v1 Announce Type: cross
Abstract: Autoregressive (AR) video diffusion enables low-latency, streamable video generation, but prediction errors often accumulate over long rollouts. Trai...
By Chenjian Gao, Zhihao Hu, Jianqi Ma, Jun Zhang, Weidong Zhang, Tianfan Xue
Block Sparse Flash Attention (BSFA) is a drop‑in replacement for FlashAttention that speeds up long‑context inference by pruning about 50% of computation and memory transfers. It selects the top‑k most important value blocks for each query using exact query‑key similarities and calibrated per‑layer, per‑head thresholds, requiring only a one‑time training‑free calibration. On Llama‑3.1‑8B, BSFA delivers up to 1.13× speedup on LongBench with a 1.1% accuracy drop and up to 1.24× on Needle‑in‑a‑Haystack retrieval with a 1% drop, while the attention kernel itself accelerates by up to 1.38×.
By Daniel Ohayon, Itay Lamprecht, Itay Hubara, Israel Cohen, Daniel Soudry, Noam Elata
OMP-MoE is a training‑free compression framework that prunes redundant experts in Mixture‑of‑Experts large language models by framing the problem as sparse signal reconstruction solved with Orthogonal Matching Pursuit. The method greedily selects expert contributions as dictionary atoms to minimize reconstruction error, then optimizes cross‑layer expert allocation via a water‑filling strategy, and finally introduces an adaptive inference mechanism (OMP‑MoE†) that dynamically adjusts expert activation based on energy prediction. Experiments on Qwen, DeepSeek‑V2, GPT‑OSS, and Mixtral MoE show consistent performance gains at 25‑50% pruning ratios, with Qwen3‑30B‑A3B retaining 93.3% of original performance at 50% compression while achieving significant speedups.
By Dezhi Li, Lujun Li, Qiyuan Zhu, Hao Gu, Bei Liu, Sirui Han, Yike Guo
The paper introduces RS-OPSD, a reliable privileged on-policy self-distillation framework designed for ultra‑high‑resolution remote sensing visual question answering. It leverages a new dataset, GeoEvidence‑6K, and a human‑feedback guided skill refinement process to provide explicit question‑relevant evidence. By incorporating context‑preserving visual privilege and correctness‑aligned distillation, RS‑OPSD achieves state‑of‑the‑art performance on several benchmarks without requiring additional visual search or tool calls at inference time.
By Chengjie Jiang, Yunqi Zhou, Jiafeng Yan, Sihang Zhao, Chun Yuan, Jing Li
LongLive‑Plug is a once‑for‑all distillation framework that learns reusable LoRA adapters on a base video diffusion model, enabling training‑free, plug‑and‑play deployment to a wide range of downstream models. These adapters provide single‑pass classifier‑free guidance, few‑step sampling, and long‑context error correction for autoregressive generation, and remain effective even when downstream models add conditioning branches or expand output channels. The authors demonstrate that the approach works on 54 downstream models across three backbone families and eight task categories, including world modeling, robotics, editing, and multimodal generation.
By Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
AdaTutoRank introduces a setwise document reranker that uses Adaptive Tutoring Optimization (ATO) to provide graded supervision across nine rubric dimensions. By generating hint‑based silver labels, reinforcement rewards, and distillation cues tailored to each rollout’s quality, the method improves credit assignment for individual documents within a set. Experiments on ten benchmarks covering Retrieval‑Augmented Generation (RAG), deep research, and setwise evaluation show that AdaTutoRank achieves superior overall performance while reducing the number of retrieval calls.
By Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu, Hongwei Zhao, Bin Li, Yu Lu, Haibo Shi
The paper investigates Neural Cellular Automata (NCAs), which are networks of recurrent cells that rely on local connectivity and asynchronous updates. It demonstrates that NCAs can solve complex visual reasoning tasks such as large mazes, Sudoku, and ARC-AGI-1, and that they generalize to larger grids, longer rollouts, and parallel trials. The study also shows that training with sample replay and stochastic perturbations enhances generalization, and that NCAs can recover from damage and scale to raw pixel reasoning.
By Mayalen Etcheverry, Pietro Miotti, Aidan Sirbu, Konstantin Sch\"urholt, Mariia Drozdova, Arna Ghosh, Blaise Ag\"uera y Arcas, James Manyika, Blake Richards, Eyvind Niklasson
arXiv:2609.36813v1 Announce Type: new
Abstract: Large language models (LLMs) exhibit strong general capabilities that mechanistic interpretability has attributed to sparse computational circuits. How...
By Chuanpu Liu, Miao Yu, Yikai Cai, Yuanhe Zhang, Zhenhong Zhou, Li Sun, Zuming Jiang, Yufei Guo
arXiv:2609.37818v1 Announce Type: cross
Abstract: Empathetic spoken dialogue requires models to use both what is said and how it is said to decide how to respond. Explicit CoT can improve paralinguis...
By Shengbo Cai, Yuxiang Wang, Jingran Xie, Zhisheng Zhang, Shun Lei, Di Cao, Teddy Sun, Zhiyong Wu
The paper introduces Q-TOFC, a query‑guided task‑oriented visual feature compression method that uses residual vector quantization to encode merged features as compact codebook index sequences. By incorporating query relevance into feature aggregation and adding a quantization error compensation adapter, Q‑TOFC reduces visual payload by 53.6% compared to previous TOFC while preserving task performance. Experiments across seven multimodal benchmarks and latency tests confirm its effectiveness under bandwidth‑constrained uplinks.
By Luning Pang, Cheng Yuan, Jiawei Shao, Mingtao Huang, Yuan Shen
arXiv:2609.37066v1 Announce Type: cross
Abstract: Post-training is central to mathematical reasoning in modern large language models (LLMs), but endpoint pass@1 alone underidentifies what has changed...
By Hongyang Li, Yiming Zhu, Xiao Li, Caesar Wu, Said Mammar, Pascal Bouvry
HyperZip introduces an efficient data compression framework that uses diffusion-based large language models (dLLMs) with Multi-Token Prediction to speed up compression. It addresses the trade‑off between throughput and compression rate by employing a hypernetwork that generates data‑specific updates from a context representation, allowing the dLLM to adapt to target data without costly fine‑tuning. Experiments show HyperZip outperforms state‑of‑the‑art baselines in both compression rate and speed.
By Thai Nguyen, Khang Tran, NhatHai Phan
The paper introduces Agent Distillation, a framework for transferring task‑solving knowledge from a teacher agent to a student agent. It categorizes where this knowledge is retained—within the model, as artifacts, through the execution harness, or across substrates—distinguishing transfer evidence from outcomes. An evaluation framework is proposed to link retention to causal contribution and practical utility, aiming to support reliable, maintainable, and safe agent development.
By Ziluowen Luo, Senzhang Wang, Chaozhuo Li, Jun Yin, Hao Yan, Ming Cheng, Chenxu Wang, Songyang Liu, Litian Zhang, Qiwei Ye, Zheng Liu, Philip S. Yu
The paper introduces FastRL, a reinforcement learning framework designed to enhance the efficiency of Group Relative Policy Optimization (GRPO) and its variants. FastRL employs an advantage-aware pruning strategy that retains high-advantage trajectories while preserving gradient diversity, and an adaptive rollout sampling mechanism that adjusts sampling scale during training based on historical pruning data. Experiments show that FastRL can be integrated into GRPO, DAPO, and GSPO, yielding a 2.07× speedup on Geometry3K and GeoQA8K-R1V and a 1.64% accuracy improvement on visual reasoning benchmarks.
By Jiahua Yang, Zhiwei Yang, Xianpeng Zhang, Dongyu Chen, Xing Chen, Tianhuang Su, Haonan Lu, Quanlong Guan, Kai Tang, Chuangchuang Wang
The paper introduces TORQUE, a framework that enhances quantization by jointly optimizing which coordinates to keep at high precision before and after applying uniform random rotations, all within a fixed bit budget. By preserving large input coordinates before rotation and the largest-magnitude coordinates after rotation, TORQUE reduces quantization error and allows efficient use of offline-optimized codebooks. The authors provide an error upper bound, prove that top‑k pre‑rotation retention is optimal for each k, and demonstrate improved accuracy‑storage tradeoffs in Gaussian models and practical tasks such as nearest‑neighbor retrieval, KV‑cache compression, and activation compression.
By Ran Ben Basat, Michael Mitzenmacher, Shay Vargaftik
ThinQuant introduces efficient rotation learning for low‑bit weight and activation quantization of large language models by reducing calibration data through a geometric selection of activations and solving a lower‑dimensional optimization problem via an ADMM algorithm. The method achieves comparable or better quantization performance with dramatically fewer calibration points, completing rotation calibration for Llama‑3‑70B in under 12 minutes and for Llama‑3.1‑405B in just over 2 hours on a single GPU. ThinQuant outperforms existing gradient‑free approaches such as DartQuant and gradient‑based SpinQuant in both speed and perplexity metrics on WikiText‑2.
By Mehdi Makni, Ryan Lucas, Rahul Mazumder
arXiv:2609.36525v1 Announce Type: cross
Abstract: Distributed inference has become an indispensable part of deploying medical models under practical latency, memory, and throughput constraints. Altho...
By Yifei Wang, Xiaohan Zhang, Youtao Ding, Tianlin Li, Xiaoyu Zhang, Yida Yang, Li Pan