Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,725 stories · RSS feed

arXiv Machine Learning
6d ago

EpiKV: Epiphany-Aware KV Cache Eviction Without the Attention Matrix

The paper introduces EpiKV, an epiphany‑aware key–value cache eviction strategy that avoids using the attention matrix. It leverages hidden‑state shifts and recent query–key relevance to rank cached tokens, matching or surpassing the performance of existing attention‑based eviction methods while remaining compatible with fast inference kernels. Experiments on multiple benchmarks show that EpiKV improves inference throughput without sacrificing accuracy.

By Steven Kolawole, Virginia Smith
arXiv AI
6d ago

SIPO: Unifying Reinforcement Learning with On-Policy Self-Distillation

SIPO (Self‑Instructing Policy Optimization) unifies reinforcement learning with on‑policy self‑distillation by using a contrastive self‑teacher to generate token‑level credit signals. The method samples multiple rollouts per prompt, pairs each with a reference answer and its mistakes, and uses the difference in teacher log‑probabilities to provide dense feedback while still respecting the overall task reward. Experiments on reasoning and code‑generation benchmarks show that SIPO outperforms both RLVR and OPSD baselines without requiring an external teacher or extra generation steps.

By Zhenrui Yue, Huimin Zeng, Yueqi Wang, Yaokun Liu, Fengran Mo, Jinghan Zhang, Mung Yao Jia, Gyuseok Lee, Yang Zhang, Na Wei, Dong Wang
arXiv AI
6d ago

IronLLM: Forging Compact Edge-Native Language Models for Real-Time Embodied Intelligence

IronLLM-0.6B is a 654‑million‑parameter language model engineered for efficient on‑device inference, featuring a hybrid attention architecture, X‑MTP multi‑token prediction, and a lightweight verification head that yields a 1.48× decoding speedup. Trained on roughly 6.2 trillion tokens with a quality‑oriented pipeline and further refined via Multi‑Domain On‑Policy Distillation, the model adopts an Instruct‑Only design to meet low‑latency requirements. A lighter variant, IronLLM‑0.6B‑Light, replaces RMSNorm with Dynamic Tanh and streamlines costly components to enhance inference and quantization efficiency, offering a strong performance‑efficiency trade‑off for resource‑constrained deployment.

By Changdi Yang, Fengquan Jiao, Haochih Lin, Haoran Yang, Jing Xiao, Liangyu Huo, Suxin Lu, Tiance Chen, Wei Liu, Yinggan Xu, Yunxiang Lu, Zai Zheng, Zhirui Xie, Zhongyang Che, Ziyan Tang, Zuoxiang Zhao, Jian Yao
arXiv Machine Learning
6d ago

Block Sparse Flash Attention

Block Sparse Flash Attention (BSFA) is a drop‑in replacement for FlashAttention that speeds up long‑context inference by pruning about 50% of computation and memory transfers. It selects the top‑k most important value blocks for each query using exact query‑key similarities and calibrated per‑layer, per‑head thresholds, requiring only a one‑time training‑free calibration. On Llama‑3.1‑8B, BSFA delivers up to 1.13× speedup on LongBench with a 1.1% accuracy drop and up to 1.24× on Needle‑in‑a‑Haystack retrieval with a 1% drop, while the attention kernel itself accelerates by up to 1.38×.

By Daniel Ohayon, Itay Lamprecht, Itay Hubara, Israel Cohen, Daniel Soudry, Noam Elata
arXiv Machine Learning
6d ago

OMP-MoE: Efficient Expert Pruning for Mixture-of-Experts LLMs via Orthogonal Matching Pursuit

OMP-MoE is a training‑free compression framework that prunes redundant experts in Mixture‑of‑Experts large language models by framing the problem as sparse signal reconstruction solved with Orthogonal Matching Pursuit. The method greedily selects expert contributions as dictionary atoms to minimize reconstruction error, then optimizes cross‑layer expert allocation via a water‑filling strategy, and finally introduces an adaptive inference mechanism (OMP‑MoE†) that dynamically adjusts expert activation based on energy prediction. Experiments on Qwen, DeepSeek‑V2, GPT‑OSS, and Mixtral MoE show consistent performance gains at 25‑50% pruning ratios, with Qwen3‑30B‑A3B retaining 93.3% of original performance at 50% compression while achieving significant speedups.

By Dezhi Li, Lujun Li, Qiyuan Zhu, Hao Gu, Bei Liu, Sirui Han, Yike Guo
arXiv Computer Vision
6d ago

RS-OPSD: Reliable Privileged On-Policy-Self-Distillation for Ultra-High-Resolution Remote Sensing VQA

The paper introduces RS-OPSD, a reliable privileged on-policy self-distillation framework designed for ultra‑high‑resolution remote sensing visual question answering. It leverages a new dataset, GeoEvidence‑6K, and a human‑feedback guided skill refinement process to provide explicit question‑relevant evidence. By incorporating context‑preserving visual privilege and correctness‑aligned distillation, RS‑OPSD achieves state‑of‑the‑art performance on several benchmarks without requiring additional visual search or tool calls at inference time.

By Chengjie Jiang, Yunqi Zhou, Jiafeng Yan, Sihang Zhao, Chun Yuan, Jing Li
arXiv Computer Vision
6d ago

LongLive-Plug: Once-for-All Distillation for Video Generation

LongLive‑Plug is a once‑for‑all distillation framework that learns reusable LoRA adapters on a base video diffusion model, enabling training‑free, plug‑and‑play deployment to a wide range of downstream models. These adapters provide single‑pass classifier‑free guidance, few‑step sampling, and long‑context error correction for autoregressive generation, and remain effective even when downstream models add conditioning branches or expand output channels. The authors demonstrate that the approach works on 54 downstream models across three backbone families and eight task categories, including world modeling, robotics, editing, and multimodal generation.

By Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
arXiv Computation and Language
6d ago

AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research

AdaTutoRank introduces a setwise document reranker that uses Adaptive Tutoring Optimization (ATO) to provide graded supervision across nine rubric dimensions. By generating hint‑based silver labels, reinforcement rewards, and distillation cues tailored to each rollout’s quality, the method improves credit assignment for individual documents within a set. Experiments on ten benchmarks covering Retrieval‑Augmented Generation (RAG), deep research, and setwise evaluation show that AdaTutoRank achieves superior overall performance while reducing the number of retrieval calls.

By Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu, Hongwei Zhao, Bin Li, Yu Lu, Haibo Shi
arXiv AI
6d ago

Reasoning with Neural Cellular Automata

The paper investigates Neural Cellular Automata (NCAs), which are networks of recurrent cells that rely on local connectivity and asynchronous updates. It demonstrates that NCAs can solve complex visual reasoning tasks such as large mazes, Sudoku, and ARC-AGI-1, and that they generalize to larger grids, longer rollouts, and parallel trials. The study also shows that training with sample replay and stochastic perturbations enhances generalization, and that NCAs can recover from damage and scale to raw pixel reasoning.

By Mayalen Etcheverry, Pietro Miotti, Aidan Sirbu, Konstantin Sch\"urholt, Mariia Drozdova, Arna Ghosh, Blaise Ag\"uera y Arcas, James Manyika, Blake Richards, Eyvind Niklasson
arXiv Computer Vision
6d ago

Task-Oriented Visual Feature Compression via Residual Vector Quantization for Device-Edge Multimodal Inference

The paper introduces Q-TOFC, a query‑guided task‑oriented visual feature compression method that uses residual vector quantization to encode merged features as compact codebook index sequences. By incorporating query relevance into feature aggregation and adding a quantization error compensation adapter, Q‑TOFC reduces visual payload by 53.6% compared to previous TOFC while preserving task performance. Experiments across seven multimodal benchmarks and latency tests confirm its effectiveness under bandwidth‑constrained uplinks.

By Luning Pang, Cheng Yuan, Jiawei Shao, Mingtao Huang, Yuan Shen
arXiv AI
6d ago

HyperZip: Efficient Data Compression through Personalized Diffusion LLMs with Hypernetworks

HyperZip introduces an efficient data compression framework that uses diffusion-based large language models (dLLMs) with Multi-Token Prediction to speed up compression. It addresses the trade‑off between throughput and compression rate by employing a hypernetwork that generates data‑specific updates from a context representation, allowing the dLLM to adapt to target data without costly fine‑tuning. Experiments show HyperZip outperforms state‑of‑the‑art baselines in both compression rate and speed.

By Thai Nguyen, Khang Tran, NhatHai Phan
arXiv AI
6d ago

Distilling Agentic Systems: A Roadmap across Models, Artifacts, and Harnesses

The paper introduces Agent Distillation, a framework for transferring task‑solving knowledge from a teacher agent to a student agent. It categorizes where this knowledge is retained—within the model, as artifacts, through the execution harness, or across substrates—distinguishing transfer evidence from outcomes. An evaluation framework is proposed to link retention to causal contribution and practical utility, aiming to support reliable, maintainable, and safe agent development.

By Ziluowen Luo, Senzhang Wang, Chaozhuo Li, Jun Yin, Hao Yan, Ming Cheng, Chenxu Wang, Songyang Liu, Litian Zhang, Qiwei Ye, Zheng Liu, Philip S. Yu
arXiv AI
6d ago

Learn from the Gap: Differential-Aware Advantage Pruning with Adaptive Rollout Sampling for GRPO

The paper introduces FastRL, a reinforcement learning framework designed to enhance the efficiency of Group Relative Policy Optimization (GRPO) and its variants. FastRL employs an advantage-aware pruning strategy that retains high-advantage trajectories while preserving gradient diversity, and an adaptive rollout sampling mechanism that adjusts sampling scale during training based on historical pruning data. Experiments show that FastRL can be integrated into GRPO, DAPO, and GSPO, yielding a 2.07× speedup on Geometry3K and GeoQA8K-R1V and a 1.64% accuracy improvement on visual reasoning benchmarks.

By Jiahua Yang, Zhiwei Yang, Xianpeng Zhang, Dongyu Chen, Xing Chen, Tianhuang Su, Haonan Lu, Quanlong Guan, Kai Tang, Chuangchuang Wang
arXiv AI
6d ago

TORQUE: Optimizing What (not) to Quantize Before and After Rotation

The paper introduces TORQUE, a framework that enhances quantization by jointly optimizing which coordinates to keep at high precision before and after applying uniform random rotations, all within a fixed bit budget. By preserving large input coordinates before rotation and the largest-magnitude coordinates after rotation, TORQUE reduces quantization error and allows efficient use of offline-optimized codebooks. The authors provide an error upper bound, prove that top‑k pre‑rotation retention is optimal for each k, and demonstrate improved accuracy‑storage tradeoffs in Gaussian models and practical tasks such as nearest‑neighbor retrieval, KV‑cache compression, and activation compression.

By Ran Ben Basat, Michael Mitzenmacher, Shay Vargaftik
arXiv AI
6d ago

ThinQuant: Scalable Rotation Learning for Weight and Activation Quantization of LLMs

ThinQuant introduces efficient rotation learning for low‑bit weight and activation quantization of large language models by reducing calibration data through a geometric selection of activations and solving a lower‑dimensional optimization problem via an ADMM algorithm. The method achieves comparable or better quantization performance with dramatically fewer calibration points, completing rotation calibration for Llama‑3‑70B in under 12 minutes and for Llama‑3.1‑405B in just over 2 hours on a single GPU. ThinQuant outperforms existing gradient‑free approaches such as DartQuant and gradient‑based SpinQuant in both speed and perplexity metrics on WikiText‑2.

By Mehdi Makni, Ryan Lucas, Rahul Mazumder