Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,914 stories · RSS feed

arXiv Machine Learning
Sep 17

The Unbearable Weight: Scaling Models and Methods for UAV Audio Classification

The paper investigates how to balance model size and fine‑tuning strategy for UAV audio classification. Using a dataset of 3,100 clips across 31 drone classes, it compares transformer and convolutional backbones under full fine‑tuning, classifier‑only fine‑tuning, and four parameter‑efficient fine‑tuning methods. Results show that selective batch‑norm tuning of EfficientNet‑B7 yields the best accuracy (97.65%) while updating less than 0.5% of parameters, and that lightweight CNNs generally outperform transformers in both accuracy and efficiency.

By Andrew P. Berg, Qian Zhang, Mia Y. Wang
arXiv Machine Learning
Sep 17

Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control

Zing-0.5 is a 5B autoregressive world model that enables users to explore and influence generated worlds through joint keyboard and online text control. It integrates unified action and text conditioning, event-scale supervision for incremental generation, and low-cost real-time interaction, achieving high scores on WBench Navigation. The authors release model weights, inference code, and a serving implementation to support further research on playable generated worlds.

By Mingyang Chen, Shengdong Chen, Xiaoxiao Fu, Bosheng Gong, Haoyuan Guo, Bowen Li, Jiawen Li, Kejun Li, Tianpeng Li, Yin Liu, Haoze Sun, Zeyang Tian, Meng Wang, Xinmiao Wu, Jiangqiao Yan, Zining Zhao
arXiv Machine Learning
Sep 17

The Operable Pareto Front: Distilling Offline Search into Run-Time Control for Multi-Objective UAV Edge-Computing Scheduling

The paper introduces PrefDT, a preference-conditioned Decision Transformer designed for multi‑objective scheduling of UAV mobile edge computing fleets. PrefDT accepts a desired energy‑delay trade‑off as input, enabling a single offline‑trained model to generate any point on the Pareto front during runtime. The authors employ attention pooling with a per‑user bypass to maintain scheduler operation when user reports are lost, and a distillation pipeline to create a preference‑labeled flight corpus, achieving superior trade‑off curves and tight energy budget adherence in simulations.

By Qiao Liao, Zhiyong Feng, Bin Wu, Guodong Fan
arXiv Machine Learning
Sep 17

A Calibrated Instrument for Measuring How Inference Optimizations Affect Output Quality

The paper introduces a calibrated instrument for rigorously measuring how inference optimizations—such as quantization, early‑exit, and speculative decoding—affect the output quality of large language models. It uses a formally calibrated LLM judge that verifies no systematic bias between statistically equivalent outputs and includes a null condition to ensure measured differences are zero. Applying this method, the authors find that a 4‑bit model is indistinguishable from its 16‑bit counterpart, while 3‑bit quantization and early‑exit techniques incur measurable quality losses that vary by language and task, and that token‑certainty‑based acceptance rules cannot reliably identify impactful errors.

By Jerry Kaplan
arXiv Machine Learning
Sep 17

Token Latency Fairness: Performance Isolation for Multi-Tenant LLM Serving

The paper introduces FairInference, a system that guarantees token-level latency isolation for well-behaved clients in multi-tenant LLM serving. It provides a δ-token fairness guarantee, ensuring that a token generated in isolation within time d will be produced within d + δ in a shared environment. The approach enforces per-token deadlines, bounds GPU compute sharing delays, and accounts for shared KV cache overhead, leading to reduced latency spikes and higher overall throughput compared to existing LLM serving systems.

By Dev Bali, Soujanya Ponnapalli, Yichuan Wang, Natacha Crooks, Scott Shenker, Matei Zaharia
arXiv Machine Learning
Sep 17

RecMorph: Topology-Guided Spatial Recurrence for Generalized Morphology Control

RecMorph introduces a topology‑guided spatial recurrent architecture for generalized morphology control, converting a kinematic tree into a sequence that enables joint cross‑limb communication and representation transformation. The design incorporates residual preservation, RMS normalization, and input‑dependent channel modulation to stabilize repeated spatial transformations, achieving linear token complexity. Across five UNIMAL tasks and a four‑platform quadruped setting, RecMorph outperforms existing controllers in training performance, inference throughput, and generalization to unseen bodies with up to 30 limbs, while also demonstrating robust real‑world performance on Go1/Go2 trials.

By Quanrui Rao, Yong Liu, Xueming Xiao, Yingbo Luo, Kun Wu, Zhenyu Xu, Meibao Yao
arXiv Machine Learning
Sep 17

TwinMark: A Unified Watermark for Provable Survival Under Feature and Logit Distillation

TwinMark is a watermarking scheme that embeds a single SHAKE128 secret into a vision model using two complementary linear functionals of model-output summaries: a covariance projector (cov‑Feat) and a class‑conditional Fisher‑aligned linear carrier (cc‑FALC). These readouts cover both classifier APIs attacked by KL knowledge distillation and representation‑only hosts attacked by feature‑matching distillation, each providing a teacher‑measurable a posteriori certificate. Across 13 attacks on datasets such as CIFAR‑10, CIFAR‑100, and Mini‑ImageNet, TwinMark remains detectable on every post‑attack model that retains task utility, survives cross‑architecture distillation onto ResNet‑18/50, VGG‑16, and MobileNet‑V3, and can be ported to GNSS few‑shot, VOC detection, ISIC segmentation, and STL‑10 SimCLR.

By Redwanul Karim, Tobias Feigl, Christopher Mutschler, Felix Ott
arXiv AI
Sep 17

How to Compress KV Cache in RL Post-Training? Shadow Mask Distillation for Memory-Efficient Alignment

The paper addresses the memory bottleneck in reinforcement learning for large language models caused by the large Key-Value (KV) cache during rollout phases. It highlights that while KV cache compression can reduce memory usage, it introduces a significant off‑policy bias that standard statistical corrections cannot adequately mitigate. The authors argue that even tiny compression errors are amplified by RL’s instability, leading to inefficient learning.

By Rui Zhu, Weiheng Bai, Qiushi Wu, Yang Ren, Haixu Tang, Yuchu Liu
arXiv Machine Learning
Sep 17

ASPIRE: Asynchronous Batched Self-Speculative Decoding for Long-Context LLM Inference

ASPIRE introduces a non‑synchronized batched self‑speculative decoding framework for long‑context LLM inference, addressing the memory bottleneck of attention by drafting tokens with sparse attention and verifying them with full attention. It combines a unified mixed forward pass, a lightweight online speculation scheduler that lets each request independently decide when to verify, and an intra‑draft refresh layer that updates the sparse context at every draft step. Experiments on three models and five benchmarks show 1.70–4.58× speedup over autoregressive baselines and a 27% average improvement over the strongest prior self‑speculative methods.

By Amir Ziashahabi, Hossein Entezari Zarch, Lei Gao, Murali Annavaram, Salman Avestimehr
arXiv Computation and Language
Sep 17

SEA-LION-v4.8: A Technical Report

The report introduces Nemotron-SEA-LION-v4.8, a family of Southeast Asian language models built on NVIDIA Nemotron 3, featuring 30B-A3B and 120B-A12B variants with both base and post‑trained checkpoints. The models are fine‑tuned on Southeast Asian, reasoning, code, and multilingual parallel datasets, then further refined with supervised fine‑tuning and online on‑policy distillation. On the SEA‑HELM benchmark, the 30B-A3B model raises the overall SEA score from 46.06 to 51.57, while the 120B-A12B model jumps from 49.30 to 63.44, with the largest improvements seen in instruction following, natural language reasoning, and understanding across seven Southeast Asian languages.

By Ahmed Mohammad Dabeer (David Wang Dawei), Ahn Jeongmi (David Wang Dawei), Anocha Sutaveephamochanon (David Wang Dawei), Antonyrex Sajeban (David Wang Dawei), Aulia Adila (David Wang Dawei), Chan Hok Teng (David Wang Dawei), Adwin (David Wang Dawei), Cheng Zi Yi (David Wang Dawei), Nicholas Zhuang Ziyi (David Wang Dawei), Choa Hsueh Mei Esther (David Wang Dawei), David Ong Tat-Wee (David Wang Dawei), Evelyn Tan Chor Phin (Li Chunren), Heng Cheng Peng (Li Chunren), Jonathan (Li Chunren), Lee Chwan Ren (Li Chunren), Leong Wai Yi (Huang Wenzong, Raymond), Leong Wei Qi (Huang Wenzong, Raymond), Leslie Teo Eng Sipp (Huang Wenzong, Raymond), Liew Rachel (Huang Wenzong, Raymond), Limkonchotiwat Peerat (Huang Wenzong, Raymond), Montalan Jann Railey Estrada (Huang Wenzong, Raymond), Muhammad Ridzuan Bin Mokhtar (Huang Wenzong, Raymond), Nagarajan Karthik (Huang Wenzong, Raymond), Ng Boon Cheong (Huang Wenzong, Raymond), Raymond (Huang Wenzong, Raymond), Ngui Jian Gang (Chen Xiaowei), Nguyen Thanh Ngan (Chen Xiaowei), Tasawong Panuthep (Chen Xiaowei), Pereira Mark Gregory (Chen Xiaowei), Phang Shi Wei Benjamin (Chen Xiaowei), Poon Yip Hung (Chen Xiaowei), Joseph (Chen Xiaowei), Rengarajan Hamsawardhini (Chen Xiaowei), Siow Wei Kang Bryan (Chen Xiaowei), Tai Ngee Chia (Chen Xiaowei), Tan Choon Meng (Chen Xiaowei), Tan Le Min (Chen Xiaowei), Sheryl (Chen Xiaowei), Tan Siao Wei (Chen Xiaowei), Tan Yi Xian, Tee Jun Yun, Teng Kok Wai, Tjhi William Chandra, Tuchinda Pume, Wu Donghang, Yong Xianbin, Yosephine, Zhang Zhou
arXiv AI
Sep 17

rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference

The paper introduces rMuscle, a real‑time Vision‑Language‑Action inference framework that mimics human muscle memory to accelerate robotic decision making. By exploiting repeated task similarity, rMuscle uses a dual‑phase cache: a Context Cache reuses visual‑token outputs and an Action Cache reuses neuron activation patterns, reducing computation and weight accesses. Experiments on RTX 4090 and Jetson Thor show 1.29–1.42× speedups on LIBERO, RoboTwin, and physical manipulation tasks while preserving success rates on real robots.

By Kaijun Zhou, Zhiyang Li, Le Chen, Jinyu Gu