Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,834 stories · RSS feed

arXiv Computation and Language
Sep 23

Semantic Self-Distillation for Language Model Uncertainty

Semantic Self-Distillation (SSD) is a method that distills the semantic dispersion of sampled answers from large language models into lightweight student models. These students estimate a prompt-conditioned density before answer generation, providing a prompt-level uncertainty signal via entropy and an answer-level reliability measure through probability density. Experiments on TriviaQA and MMLU show that SSD matches the teacher’s uncertainty estimates while enabling additional tasks such as hallucination prediction, out-of-domain detection, and multiple-choice answer selection.

By Edward Phillips, Sean Wu, Fredrik K. Gustafsson, Boyan Gao, David A. Clifton
arXiv Machine Learning
Sep 23

CompKV: Compensation-Aware KV Selection for Long-Context LLM Inference

CompKV introduces a compensation‑aware sparse attention framework for long‑context LLM inference. It partitions tokens into blocks and optimizes token selection to minimize the error introduced by block‑level mean compensation, using compact block‑level statistics. Experiments on RULER and LongBench‑Pro show CompKV outperforms other sparse baselines and achieves up to a 6.85× speedup over full attention.

By Zhen Huang, Ruizhe Yao, Danyi Liu, Xinrui Chen, Shuwei Li, Siru Zhong, Zijian Cao, Yushan Lai, Mingming Guo, Weijie Zheng, Haohuan Fu
arXiv Machine Learning
Sep 23

Geometry-Aware Hyperbolic Residual Quantization

The paper introduces Geometry‑Aware Hyperbolic Residual Quantization, a method that adapts residual vector quantization to hyperbolic space while preserving its telescoping structure. It achieves this by using Hyperbolic Residual Aggregation in the forward pass and a discounted Hyperbolic Straight‑Through Estimator in the backward pass, thereby avoiding geometric inconsistencies and unstable gradients. Experiments on hierarchical prediction, recommendation, image tokenization, and neural audio coding demonstrate improved stability and structural organization of hyperbolic residual codes, with a noted trade‑off between compression efficiency and hierarchical organization.

By Alessio Colombo, Melika Ayoughi
arXiv Machine Learning
Sep 23

One-Step Generative Surrogate Models via Block-Triangular Joint Drifting

The paper introduces block‑triangular joint drifting, a method that applies a projected drift field to the joint distribution of consecutive states, enabling one‑step generative surrogate models for stochastic transition dynamics. This architecture preserves the current‑state marginal while directly sampling the conditional distribution of next states, allowing stochastic trajectories to be generated with a single model evaluation per time step. Experiments show that the approach achieves accurate marginal and trajectory‑dependent statistics with favorable accuracy‑cost tradeoffs compared to deterministic, diffusion, flow, and distillation‑based generative surrogates.

By Nicholas Geissler, Shreya Jha, Ricardo Baptista, Benjamin Peherstorfer
arXiv Machine Learning
Sep 23

Not All 4-bit Quantizers Are Equal: Deployment-Time Mitigation of PII Leakage in Fine-Tuned Small Language Models

The paper investigates how different 4‑bit quantization techniques affect privacy when fine‑tuned small language models are deployed. It finds that methods using a calibration corpus, such as Activation‑aware Weight Quantization (AWQ) and Gradient‑based Post‑Training Quantization (GPTQ), prevent the reproduction of planted private records, whereas a calibration‑free format (GGUF Q4_K_M) leaks 5.3% of them. Across models ranging from 0.5 to 7 billion parameters, AWQ consistently leaks the least while maintaining minimal accuracy loss, indicating that the choice of 4‑bit method is a privacy decision as well as a performance one.

By Cristhian Kapelinski, Diego Kreutz
arXiv Machine Learning
Sep 23

What Does Chain-of-Thought Entropy Measure? A Channel Audit of Scaffolding, Routing, and Content

The paper investigates how entropy over chain‑of‑thought tokens influences policy decisions such as gradient application, pruning, and collapse detection. By separating scaffold tokens from substantive content, the authors analyze entropy, Kullback–Leibler divergence, and entropy velocity for each channel, proving differences between raw and content conventions and bounding answer diversity. Empirical results across 23 configurations show that scaffold tokens can account for up to 41% of high‑entropy positions, with entropy share growing through distillation, while content conventions outperform raw surprisal on compression tasks and reveal significant answer leakage in re‑fed chains.

By Marios Papamichalis, Regina Ruane
arXiv Machine Learning
Sep 23

MIND the Gap: A Geographic Implicit Neural Representation with Adjustable Spatial Scale

The paper introduces MIND, a method that distills specialist geospatial model embeddings into a single generalist coordinate embedding with adjustable spatial granularity, using nested supervision across multiple embedding dimensions. MIND’s design allows downstream predictors to use only leading chunks or apply a Chunked Penalty to downweight finer details without retraining the INR. The authors evaluate MIND on CoordBench, a large INR benchmark of 52 datasets and 78 targets, and report that MIND and its Chunked Penalty variant achieve the highest regression and classification scores, especially under regional holdout, establishing a new state‑of‑the‑art for geographic implicit neural representations.

By Isaac Corley, Arjun Rao, Esther Rolf, Konstantin Klemmer, Evan Shelhamer, Nils Lehmann, Marc Ru{\ss}wurm, Gengchen Mai, Nathan Jacobs, Hannah Kerner
arXiv Machine Learning
Sep 23

RootQuantV2: Adapting a Vision Foundation Model for Root-Trait Regression from Minirhizotron Imagery

RootQuantV2 adapts a frozen DINOv3 Vision Transformer to predict root length and surface area directly from minirhizotron images, eliminating the need for manually traced masks. By training only 11.9 M parameters (3.78 % of the model), it achieves R² values of 0.950 for length and 0.930 for area, improving RMSE by 24.3 % and 20.7 % over the previous RootQuant CNN approach. The method repurposes existing numeric archives of root traits for high‑throughput, automated phenotyping in field‑grown crops.

By Kinjalk Parth, Sebastian Varela, Andrew D. B. Leakey
arXiv Machine Learning
Sep 23

Ladders of Thought: A Self-Evolving Curriculum of Progressively Simplified Reasoning Traces

Ladders-of-Thought (LoT) is a framework that enhances reasoning in small- to mid-scale large language models by automatically generating easier variants of reasoning problems and organizing them into difficulty buckets. It uses a self‑evolving bandit scheduler to adaptively allocate training, improving performance across math and multi‑hop reasoning tasks on 1–8 B models. LoT achieves significant gains (e.g., +32 pp on AddSub, +16 pp on QASC) and converges faster than staged curricula.

By Minghui Liu, Thomas Magelinski, Dehao Yuan, Qi Yu, Furong Huang
arXiv Machine Learning
Sep 23

Magnitude Profile Pruning: Calibration-Free Structured Attention Head Removal for Transformer Compression

Magnitude Profile Pruning introduces a training‑free, calibration‑free method for removing attention heads in Transformer models by statistically detecting outliers in weight row norms. Heads whose projection weights fall within the bulk of the distribution are pruned, while outlier heads are retained. Across several models, the MP‑G variant achieves superior perplexity at various sparsity levels and yields significant parameter and FLOP reductions without requiring forward passes, calibration data, or gradient computations.

By Kasun Dewage, Marianna Pensky, Heranga K. Rathnasekara, Suranadi De Silva
arXiv Machine Learning
Sep 23

Damage Predicts Recovery: When Calibration Data Matters in Compressing Financial LLMs

The paper investigates whether domain‑matched calibration data is necessary when compressing large language models for financial tasks. It finds that if compression causes little task damage, the choice of calibration corpus has minimal impact, whereas significant damage—especially from pruning—can be mitigated by using a finance‑specific calibration set (FinMix). The study tests this across multiple models, compression settings, and financial tasks, consistently supporting the link between damage and recovery.

By Junyi Ye, Mengjia Yu, Debapriya Hazra, Guiling Wang
arXiv Machine Learning
Sep 23

PP-Net: A Hybrid Physical-Prior Neural Network for Scattered Light Removal in Biomedical Images on Embedded Devices

PP‑Net is a hybrid physical‑prior neural network designed to remove scattered light from biomedical images on embedded devices. It combines a denoising network (DFN‑Net), a scattering‑map estimator (ASAP), and a refinement network (GF‑Net) to fuse a physics‑based prior with denoised observations. The method uses progressive synthetic training and cross‑domain transfer to reduce reliance on paired ground truth, achieving significant PSNR, SSIM, and NIQE improvements over baselines while maintaining an inference latency of about 200 ms per 512×512 image on edge hardware.

By Yongfei Guo, Tingjin Chu, Mengzhuo Liu, Hongwei Lou, Yuanhao Gong
arXiv Machine Learning
Sep 23

GTR: Gated Token Recurrence for Efficient Dense Prediction

The paper introduces Gated Token Recurrence (GTR), a softmax‑free recurrent vision backbone that replaces global softmax attention with gated linear attention, alternating scan directions, and enhanced SwiGLU blocks. GTR is distilled from a DINOv3 teacher using only final‑layer patch‑token alignment, and achieves strong performance on COCO object detection (58.9 box AP) with very low latency (1.908 ms on an RTX 4090). The backbone also transfers to multiple dense prediction tasks and runs efficiently on edge hardware via a specialized CUDA operator and TensorRT deployment.

By Zhe Feng, Longfei Liu, Wei Liu, Kai Chen, Jiangjiang Kong, Wei Zhou, Yifeng Qian, Dexiong Chen, Xuanlong Yu, Xi Shen
arXiv Machine Learning
Sep 23

Hierarchical Sparse Bayesian Multitask Learning for Disease Prediction in Pooled Microbiome Studies

This paper introduces a hierarchical Bayesian multitask learning model that assumes a shared sparsity structure across different binary classification tasks. The authors develop a variational inference algorithm for efficient posterior approximation and evaluate the method on synthetic data and pooled microbiome studies. Results show superior support recovery in synthetic experiments and robust, well‑calibrated predictions with informative taxa selection in microbiome classification.

By Haonan Zhu, Andre R. Goncalves, Camilo Valdes, Hiranmayi Ranganathan, Boya Zhang, Jose Manuel Mart\'i, Car Reen Kok, Monica K. Borucki, Nisha J. Mulakken, James B. Thissen, Crystal Jaing, Alfred Hero, Nicholas A. Be