Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,731 stories · RSS feed

arXiv AI
Sep 28

TISD: On-Policy Self-Distillation with Trajectory Intervention

The paper introduces TISD, a trajectory-intervention self-distillation method that forces a teacher-selected branch action and then lets the student generate the suffix, distilling the full trajectory under a privileged-context-conditioned teacher. This approach addresses a data-collection bottleneck in on‑policy self‑distillation by exposing successor contexts that the student would otherwise miss. Experiments on coding and science domains show modest but consistent improvements in average performance metrics compared to baseline methods.

By Taeckyung Lee, Rinat Amankos, Jeonghye Kim, Hyungjun Yoon, Woogyeol Jin, Sung-Ju Lee
arXiv AI
Sep 28

Momentum-Guided Federated Split Distillation for Personalized Temporal Edge Intelligence

The paper introduces a momentum-guided federated split distillation framework for personalized temporal edge intelligence. It presents TeRR-SAtt, a temporal reservoir student attention design that uses fixed reservoir representations, a lightweight temporal student, and personalized output modules. Additionally, it proposes AMGF, an anticipatory momentum-guided fusion mechanism that clusters clients via learning momentum and generates specialized teacher updates. Experiments on real-world smart‑building data show that TeRR-SAtt cuts edge training latency by 65.50%, inference latency by 44.70%, training memory usage by 18.40%, and inference CPU usage by 33.10% compared to baselines, while AMGF improves local learning by up to 35.31% in RMSE over global updates.

By Ahmed-Rafik Baahmed (LINEACT), Jean-Fran\c{c}ois Dollinger (LINEACT), Amine Brahmia (LINEACT), Mourad Zghal (LINEACT)
arXiv AI
Sep 28

G2MAF: Test-Time Gradient Guidance for Multi-Agent Flow Policies

G2MAF is a test‑time refinement framework for offline multi‑agent reinforcement learning that applies a single globally normalized, projected critic gradient to adjust all agents’ actions while keeping them close to a frozen policy proposal. The method improves performance on 24 Multi‑Party Environment (MPE) and StarCraft Multi‑Agent Challenge (SMAC) benchmarks, achieving mean relative gains of 9.2% on MPE and 8.9% on SMAC, with only a 6% increase in inference latency.

By Guowei Zou, Haitao Wang, Guoxin Wang, Zhiquan Chen, Beiwen Zhang, Guojie Wang, Hejun Wu
arXiv AI
Sep 28

The Right Information Extraction Pipeline Depends on the Document: Accuracy-Energy Trade-offs for Small, Local Models

The study investigates how the choice between processing page images or parsed text in an information extraction pipeline depends on the document’s layout, focusing on privacy‑sensitive, on‑premise scenarios with small models (≤8 B parameters). It evaluates accuracy and energy consumption across input representations, model families, and inference settings, finding that batching dramatically reduces energy use, FP8 quantization offers modest savings, and neural OCR is far more energy‑intensive than classical OCR. The optimal representation varies: vision‑language models excel on layout‑rich documents, while small text‑only models with a cheap parser perform best on near‑plain‑text contracts, achieving higher accuracy and lower energy than any vision‑language setup.

By Christoph Walser, Mauricio Fadel Argerich, Jonathan F\"urst
arXiv AI
Sep 28

Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents

The paper introduces LOHA, a context layout that compresses older tool observations into soft tokens while keeping the agent’s own turns and the last K observations in plain text, and ACD, a training method that distills full‑text predictions into this latent representation while anchoring behavior on plain text. This approach reduces context per call by up to 57% without significant loss in resolve rates, and improves instance throughput in single‑GPU serving. Experiments on SWE‑bench Verified show that K=3 yields a 43–57% compression with only modest performance impact, while larger windows favor task performance over compression.

By Zhensheng Zou (Peking University), Guoqing Wang (Peking University), Dan Hao (Peking University)
arXiv AI
Sep 28

Prompt Minimization: Reducing Input Redundancy Without Sacrificing Output Fidelity

The paper investigates prompt minimization, aiming to reduce prompts to their smallest, most information-dense form without losing output fidelity. It argues that shorter prompts lower computational overhead and inference latency, especially when large contexts are unnecessarily included, and that longer prompts can harm LLM reasoning and accuracy. The authors propose three frameworks to identify minimal prompts and show that these often produce outputs comparable to longer versions, highlighting redundancy in the input space and opening new avenues for efficient prompt engineering.

By Marius F. R. Juston, Kevin A. Karim, Jonathan Gao, Kevin C. Li, Rudhi Bashambu
arXiv AI
Sep 28

DeepEdu-v1: Efficient and Scalable Agentic LLMs for Vietnamese Education

DeepEdu‑v1 is an AI‑tutoring system tailored for Vietnamese education that addresses data‑sovereignty and local curriculum alignment issues. It uses a long‑context inference engine to reduce retrieval calls and prefill latency by about 35%, and a self‑improving agentic layer that curates verified local knowledge without fine‑tuning. In deployment, DeepEdu achieves nearly twice the speed of standard vLLM serving and raises agentic accuracy from 70.0% to 79.5% on complex tasks, especially in financial reasoning and interactive‑agent benchmarks.

By Quang Nguyen, Hieu Nguyen, Hien Hoang, Toan Pham, Cong Tran, Nam Vu
arXiv AI
Sep 28

Adaptive multi-resolution Gaussian processes: Scalable exact inference with naturally data-sparse covariance matrices

The paper introduces an adaptive multi‑resolution Gaussian process framework that achieves scalable, exact inference by constructing a naturally data‑sparse covariance matrix using basis functions anchored directly to samples. By shrinking the support domains of these basis functions, the resulting matrix has limited block sizes, ensuring sparsity and enabling efficient computation of its inverse via a sparse Cholesky algorithm. The authors demonstrate that this approach yields exact inference with training cost ≠≠ O(n log^2 n) and prediction cost ≠≠ O(log^d n), while also improving predictive uncertainties through an augmented basis function.

By Yanchuang Cao, Jun Liu, Tengchao Yu, Heng Yong
arXiv AI
Sep 28

PolicyAttention: Softmax Attention Implements Policy Mirror Descent for Closed-Loop Control

The paper investigates whether causal softmax attention can realize policy mirror descent (PMD) as a repeated controller rather than a one‑step algebraic identity. It constructs a fixed causal‑softmax actor–environment–one‑step‑critic protocol, detailing actor, routing, sampling, and normalization residuals, and shows that a frozen one‑step audit model closely approximates PMD. Empirical results demonstrate that the learned actor with an exact one‑step critic achieves median policy loss only about 5% higher than the exact PMD oracle across multiple control settings.

By Yuhe Sui, Yingzhi Tang, Shufang Chen
arXiv AI
Sep 28

MedTokenBudget: Lesion-Preserving Token Routing for Dermoscopic Image Classification

MedTokenBudget introduces a supervised token routing framework for Vision Transformers applied to dermoscopic image classification. Its Lesion-Aware Token Scoring (LATS) module combines attention entropy, feature norm, and local feature contrast to select the top‑K patches under a target budget, trained with curriculum learning, diversity regularization, attention distillation, and lesion‑mask supervision. On the ISIC 2019 dataset, mask‑supervised LATS outperforms Random and ToMe at headline budgets while retaining more lesion patches.

By Zhexiang Li
arXiv AI
Sep 28

MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation

MOPD‑Router rethinks teacher routing in multi‑teacher on‑policy distillation by routing supervision over the full teacher pool at each token, eliminating the need for prompt‑level domain labels or a separate routing model. The framework offers a plug‑in interface for various metrics, and introduces ExpertAlign, which scores teachers based on how well their corrections reflect their specialized post‑training knowledge. Experiments on both unlabeled and domain‑labeled mixtures show that ExpertAlign outperforms existing methods, improving overall scores by up to 12.3% on unlabeled data and 7.8% on domain‑labeled data. whyItMatters":"Token‑level routing enables the use of complementary supervision across domains without relying on domain labels, leading to significant performance gains in multi‑teacher distillation settings."

By Tianze Xu, Yanzhao Zheng, Zhentao Zhang, Yuanqiang Yu, Chao Ma, Jihuai Zhu, Lelun Wu, Lyumanshan Ye, Pengfei Liu, Baohua Dong, Hangcheng Zhu, Ruohui Huang, Gang Yu
arXiv AI
Sep 28

Persistent Negatives for Adversarial Black-Box On-Policy Distillation

The paper introduces Persistent Negative Adversarial Distillation, a method that improves black-box on-policy distillation by maintaining a live pool of historical teacher–student comparisons to stabilize the discriminator’s negative distribution. By anchoring the discriminator with these persistent negatives, the approach reduces reward-estimation error and yields smoother, higher-performing student policies across multiple benchmarks. The study demonstrates that the choice of negative samples is a critical design factor in effective black-box distillation.

By Haixu Ma, Saad Lahrichi, Weiwei Li, Kevin Han, Weiqiang Wu, Peggy Yang, Dongzhuo Li, Ruiyi Li, Serena Li, Gedi Zhou, Mingze Gao, Abhishek Kumar, Xiangjun Fan, Lizhu Zhang
arXiv AI
Sep 28

G$^2$PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation

G$^2$PTQ is a post‑training quantization framework that improves large language models by combining first‑ and second‑order information in a globally supervised, block‑wise optimization. It refreshes gradient and Hessian estimates before each Transformer block and uses a trust‑region scaling mechanism to stabilize gradient steps, preventing exploding weight updates. The method achieves better alignment with full‑precision models and outperforms state‑of‑the‑art baselines across various model families and bit‑widths.

By Ruikang Liu, Haoli Bai, Yuxuan Sun, Qian Zhang, Wenzheng Cai, Yanqi Hao, Feiyu Wang, Weidong Zhong, Zhuang Wang, Tong Yang, Xiangsheng Zhou
arXiv AI
Sep 28

Teacher-Anchored Selection of Post-Training Quantized Models under Domain Shift

The paper investigates how to choose the best quantized model from a family of compressed versions when target labels are scarce or unavailable. It finds that a simple rule based on minimum teacher distortion consistently selects the same eight‑bit, per‑channel, unclipped configuration, though this does not minimize empirical target cross‑entropy. The study also shows that confidence‑based estimators perform poorly in overconfident regimes, while output‑distribution estimators can outperform the teacher in some architectures, and that combining distortion with a supervised term can improve selection. Across 134 candidate families, teacher‑anchored selection reduces mean regret with very few labels, though the benefit diminishes after about 25 labels.

By Alejandro Rodriguez Dominguez, Muhammad Shahzad, Xia Hong
arXiv AI
Sep 28

Acoustic-to-Text KV Compression for Full-Duplex Speech Models

The paper introduces acoustic-to-text KV compression for full‑duplex speech models, converting acoustic key‑value states into compact textual memory during listening‑time slack. When the KV cache exceeds a target budget, older acoustic states are evicted while transcripts and recent acoustic context are retained. Experiments on ten‑minute LongSpeech sessions show a 64.6% reduction in peak streaming KV‑cache size and improved transcription, temporal question answering, and summarization, with comparable pause‑handling, turn‑taking, and interruption performance in Full‑Duplex‑Bench.

By Yejin Lee, Seungbeom Kim, Yongha Lee, Kyuhong Shim
arXiv AI
Sep 28

Softmax Reparameterization for Output-Head Quantization

The paper introduces a post‑training softmax reparameterization technique that selects a functionally equivalent output head before quantization. By subtracting a scalar multiple of the vocabulary‑row mean from each output row and tuning this coefficient via validation KL, the method preserves the full‑precision softmax distribution while enabling efficient W4 quantization. Experiments on seven heads show significant error reductions and latency improvements, with the approach remaining complementary to other quantization strategies and transferable across datasets.

By Asim Kadav, Christian Flores, Chirag Arora, Varun Kotte, Hongbo Zheng, Lan Yan, Priya Shanmugasundaram, Tracy Holloway King
arXiv AI
Sep 28

UniAR: A Unified Framework for Autism Recognition Enhanced by Multi-View Prompt Learning

UniAR is a unified framework that improves autism spectrum disorder (ASD) recognition by using multi-granularity prompt learning and a large multimodal model to generate diagnostic descriptions at word, phrase, and sentence levels. It aligns these semantic representations with visual evidence through a Mixture-of-Experts-based Multi-Scale Alignment Module, enabling robust ASD detection across heterogeneous data types. Experiments on four brain MRI and facial expression benchmarks show that UniAR outperforms state‑of‑the‑art methods, achieving 75.9% accuracy on MRI and 91.6% on facial benchmarks, with gains of 1.5 and 1.2 percentage points respectively.

By Lei Xin, Zeheng Wang, Jiayin Zhu, Shihong Huang, Fanhu Zeng, Changjiang Jiang, Dengbo He, Yutao Yue, Zhenglun Kong
arXiv AI
Sep 28

DyMD: Preserving Interaction Dynamics through Distribution Matching Distillation in Few-Step Video World Models

DyMD introduces a Distribution Matching Distillation framework that adapts teacher supervision and critic fitting to preserve interaction dynamics in few-step video generation. By employing temporal affinity–conditioned re‑noise sampling and dynamics‑guided fake‑score tracking, DyMD balances motion recovery with visual quality. The method distills a 14B teacher into a 1.3B student that achieves significant gains on embodied‑video benchmarks and downstream action planning tasks.

By Haojun Xu, Jie Huang, Xin Lu, Mingchen Zhong, Zihao Fan, Linjiang Huang, Si Liu
arXiv AI
Sep 28

Intent2Tc: Automated Intent-to-Traffic Control Translation with Language Models

Intent2Tc is a closed‑loop, language‑model‑driven framework that translates high‑level business traffic‑shaping intents into executable Linux traffic‑control (tc) configurations. It uses an AQM‑based digital twin semantic model, automated metadata extraction, critique‑driven refinement, and Retrieval‑Augmented Generation to improve semantic consistency and configuration reliability. Evaluation on 100 RFC 9315‑compliant intents shows high semantic fidelity and deployment readiness, with Claude Sonnet‑4.6 achieving 0.98 semantic similarity and 0.045 normalized edit distance, while RAG reduces token consumption and latency for compact models.

By Andrea Masini, Sudipta Acharya, Paolo Bellavista, Luca Foschini, Burak Kantarci
arXiv AI
Sep 28

Stepwise Intrinsic Rewards for Reasoning in Large Language Models

The paper introduces Stepwise Marginal Information Gain (MIG), an intrinsic process reward that evaluates how each reasoning step of a large language model (LLM) or vision-language model (VLM) improves the likelihood of the reference answer. MIG rewards only new likelihood maxima, preventing duplicate credit, and is combined with outcome, format, and self‑distillation objectives to guide training. Experiments on eight benchmarks show that this method outperforms outcome‑only reinforcement learning and improves accuracy by up to 4.8 points over binary‑reward training, including a 12.6‑point gain on MathVerse and a 12.9‑point advantage on vision‑language transfer at 7B parameters.

By Xiangwei Wang, Wei Wang, Ken Chen, Nanduni Nimalsiri, Sachith Seneviratne, Saman Halgamuge