Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,914 stories · RSS feed

arXiv Computer Vision
Sep 17

StrucPhysVideo: Learning Physical Dynamics from Structured Captions and Robot Actions

arXiv:2609.18430v1 Announce Type: new Abstract: Modeling physical dynamics, including how objects move, interact, and change state, is central to video world models for embodied AI. We present StrucP...

By WM Team, Enhui Ma, Kaiwen Guo, Tingrui Zhang, Wei Song, Yingshui Tan, Jianhua Xu, Tong Zhang, Kaicheng Yu
arXiv AI
Sep 17

OBC-Prune: Outcome-Based Calibration for Large Reasoning Model Pruning

OBC‑Prune introduces an outcome‑based calibration approach for pruning large reasoning models, focusing on the causal importance of each reasoning sentence rather than uniform activation salience. By pairing correct and incorrect rollouts and using intervention‑based analysis, it assigns per‑token weights that guide one‑shot pruning methods such as SparseGPT, Wanda, and ALPS. Experiments on DeepSeek‑R1‑Distill‑Qwen models show consistent accuracy gains and shorter reasoning traces across multiple benchmarks at 40‑50% sparsity.

By Ha Lan Nguyen, Huy Hoang Tran, Trac-Duy Tran, Dung D. Le
arXiv AI
Sep 17

Diff-SPORT: Diffusion-based Sensor Placement Optimization and Reconstruction of Turbulent flows in urban environments

Diff-SPORT is a diffusion-based framework that integrates a generative diffusion prior, maximum a posteriori inference, and Shapley-value attribution to achieve high-fidelity reconstruction of turbulent flows and optimal sensor placement in urban environments. By training the diffusion prior once for a domain, it enables non-linear sensor placement and near-real-time flow reconstruction from sparse measurements, outperforming state-of-the-art methods and running orders of magnitude faster than RANS or LES simulations. The approach also generalizes to experimental passive scalar concentration data, demonstrating up to 57% lower reconstruction error than random sensor placement under extreme sparsity and providing compact, physically interpretable sensor configurations.

By Abhijeet Vishwasrao, Sai Bharath Chandra Gutha, Andres Cremades, Klas Wijk, Aakash Patil, H. D. Lim, Christina Vanderwel, Catherine Gorle, Beverley J McKeon, Hossein Azizpour, Ricardo Vinuesa
arXiv Machine Learning
Sep 17

FairLRF: Achieving Fairness through Sparse Low Rank Factorization

FairLRF proposes a fairness-oriented low rank factorization framework that uses singular value decomposition (SVD) to improve deep learning model fairness. By selectively removing bias-inducing elements from the unitary matrices obtained via SVD, the method reduces group disparities while preserving accuracy. Experiments demonstrate that FairLRF outperforms existing low rank factorization and state-of-the-art fairness techniques, and an ablation study explores the impact of key hyper-parameters.

By Yuanbo Guo, Jun Xia, Yiyu Shi
arXiv AI
Sep 17

ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models

ActionPiece rethinks how actions are tokenized for autoregressive vision‑language‑action models by introducing physical rank consistency (PRC) to evaluate relational fidelity of reconstructed actions. The method jointly supervises representation learning and quantization to preserve local physical distance rankings, improving both PRC and policy success. Experiments on LIBERO, LIBERO‑Plus, SimplerEnv, and VLA‑Arena show significant gains over baseline tokenizers.

By Shijie Lian, Bin Yu, Zhaolong Shen, Xiaopeng Lin, Yichao Du, Zhirui Zhang, Laurence T. Yang, Kai Chen
arXiv AI
Sep 17

Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It

The paper argues that agentic systems waste time and memory by guessing how long tool calls will take, rather than using explicit progress signals from the tools themselves. It demonstrates that tools can report their remaining work or imminent completion, and that incorporating this feedback into serving systems dramatically improves cache decisions and reduces token latency. The authors show that this approach outperforms existing predictors and works robustly across different environments.

By Yipeng Liu, Yingqiang Zhang, Feifei Li, Huanchen Zhang
arXiv AI
Sep 17

Taming the Agentic RAN: Stability-Guaranteed Arbitration of Autonomous AI Agents in O-RAN

The paper demonstrates that autonomous AI agents (rApps) in an O-RAN control plane can independently close control loops over shared radio resources, leading to unsafe interactions when multiple agents pursue different objectives. The authors introduce AURA, a lightweight arbitration layer that enforces feasibility invariants, dwell times, and a deadband to ensure stable operation. Implemented on an OpenAirInterface testbed, AURA reduces shared-state excursions by more than an order of magnitude and virtually eliminates cross-slice throughput starvation while maintaining latency compliance for the protected slice.

By Seyed Bagher Hashemi Natanzi, Bo Tang
arXiv AI
Sep 17

A Zeroth-Order Paradigm for LLM Preference Alignment

The paper introduces Comparison-based Preference Optimization (ComPO), a zeroth-order method that aligns large language models with human preferences using comparison oracles instead of direct differentiable loss optimization. It provides theoretical convergence guarantees for both offline and online variants under smoothness, gradient sparsity, and oracle compatibility assumptions, and establishes performance bounds under local coverage and in-distribution reward accuracy. Experiments on several LLMs (Mistral, Llama, Gemma-2, Qwen3, Gemma-3) show that ComPO outperforms existing direct alignment methods, achieving higher length-controlled win rates and diagnostics that suggest mitigation of likelihood displacement.

By Peter Chen, Xi Chen, Wotao Yin, Tianyi Lin
arXiv AI
Sep 17

Calendar-Structured Sparse Principal Component Analysis for Interpretable Multi-Periodic Electricity Consumption Profiles

Calendar-Structured Sparse Principal Component Analysis (Calendar-SPCA) is a new method that learns low-dimensional representations of long-term electricity consumption data by explicitly incorporating daily, weekly, and annual calendar cycles. It uses an L1 penalty and graph total variation to produce sparse, locally coherent, and directly interpretable latent factors. In experiments on the GoiEner and Low Carbon London smart‑meter datasets, Calendar-SPCA retains most of the variance of standard PCA while achieving high sparsity and clear calendar‑aligned structures.

By Carlos Quesada-Granja, Tony Castillo-Calzadilla, Carlos Rizo-Maestre
arXiv AI
Sep 17

The Inference Engineering Pareto Atlas: Which Optimizations Dominate the Cost, Quality, and Latency Frontier?

The paper presents a Pareto atlas of LLM inference optimizations, mapping cost, quality, and latency trade‑offs for Qwen2.5‑7B‑Instruct on L4, A100, and H100 GPUs. Using 54 measured configurations and a calibrated simulator, it identifies 18 of 36 setups on the Pareto frontier, showing that combined methods outperform single ones. Quality tests reveal that AWQ 4bit and FP8 weights offer significant latency reductions while largely preserving accuracy, but naive FP8 KV caching fails to answer any questions correctly.

By Srikanta Datta Tumkur, Jay Iyer, Mehar Simhadri, Sai Pavan Kumar, Sai Kapil Kumar, Ramesh Nampelly
arXiv Machine Learning
Sep 17

Long-Context Demonstration Selection Using State Space Models

The paper addresses the challenge of selecting demonstrations for long-context language model queries, where transformer inference costs grow quadratically with sequence length. It proposes two algorithms that distill transformer behavior into state space models (SSMs) with linear inference time, partitioning transformer layers into groups and estimating separate SSMs for each. The distilled SSMs achieve less than 0.7% approximation error, and in downstream tasks they reduce FLOPs by 14.2× while improving accuracy by 6.48% compared to baseline methods.

By Ziniu Zhang, Zhenshuo Zhang, Ruoxuan Xiong, Gene Cooperman, Hongyang R. Zhang
arXiv Machine Learning
Sep 17

Colla-Q: Toward Collaborative Experts in MoE Quantization via Minimax Precision Balancing

The paper introduces Colla-Q, a Mixture-of-Experts (MoE) quantization technique that uses activation entropy to allocate bit-widths across experts. By balancing performance among experts, Colla-Q improves overall MoE accuracy and reduces reliance on calibration datasets. The method aims to maintain robustness and stability in quantized MoE models.

By Eunju Shin, Jongbin Ryu
arXiv Machine Learning
Sep 17

GroupKV: Hierarchical KV Cache Management for Long-Context Diffusion LLM Inference

GroupKV is a lightweight hierarchical KV cache management system designed for long‑context diffusion large language model (dLLM) inference. It partitions the context into contiguous groups and uses coarse‑to‑fine sparse selection, cross‑layer consistency for predictive prefetching, and a staleness correction mechanism to keep the cache coherent amid dynamic KV updates. The approach also incorporates streaming prefill to lower peak memory usage, achieving up to 48× longer serviceable context, 3.73× faster inference in offload‑based settings, and competitive task accuracy.

By Jinhao Wang, Zhexin Hu, Kangjie Zhou, Xin Zhou, Fangfang Liu