Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,834 stories · RSS feed

arXiv Machine Learning
Sep 17

ADAPT: Lightweight, Long-Range Machine Learning Force Fields Without Graphs

The paper introduces ADAPT, a lightweight machine‑learning force field that replaces graph neural networks with a direct coordinates‑in‑space Transformer encoder to model all pairwise atomic interactions. Applied to silicon point defects, ADAPT reduces force prediction error by about 22% and energy prediction error by roughly 40% compared to a state‑of‑the‑art GNN model, while also cutting computational cost. This approach addresses common GNN issues such as oversmoothing, oversquashing, and poor long‑range interaction representation, which are especially problematic for point defect modeling.

By Evan Dramko, Yihuang Xiong, Yizhi Zhu, Geoffroy Hautier, Thomas Reps, Christopher Jermaine, Anastasios Kyrillidis
arXiv AI
Sep 17

Imitation Learning for Autonomous Driving in CARLA

The paper presents a compact multimodal policy trained via behavioral cloning to drive autonomously in the CARLA simulator. Using five‑frame histories of RGB images, LiDAR, telemetry, and lane waypoints, the 1.36‑million‑parameter model predicts throttle, brake, and steering at 20 Hz. Trained on 236,882 windows (≈3.3 hours of driving) from 448 captures, the policy drives for hours on both training and unseen routes without collisions, demonstrating qualitative transfer and recovery from large trajectory deviations.

By Jordy Kieto
arXiv AI
Sep 17

Pay Only for Disagreement: Certified No-Regression Verdicts for Model Updates with Matching Label-Complexity Bounds

The paper introduces DISCERN, a two-tier protocol for certifying that updates to production models do not increase risk. It first uses unlabeled data to detect benign updates based on disagreement rates, then selectively labels only disagreements through an anytime-valid confidence sequence. The method achieves finite-sample validity with label-complexity bounds of order ρ²/ε², demonstrating significant label savings and strong empirical performance across 14,000+ audit streams.

By Vishnu Bindu Balachandran
arXiv AI
Sep 17

CPR: Combining global composing, local performing and full-sequence refining in piano rendering with continuous autoregressive modelling

The paper introduces CPR, a piano rendering framework that combines continuous autoregressive modeling with local flow matching and full‑sequence refinement. It predicts continuous hidden states, generates 24 kHz acoustic latents, and upsamples to 48 kHz, while new techniques BREPA and MT‑RoPE enhance musical semantics and cross‑modal alignment.

By Chong Jing, Junan Zhang, Zhizheng Wu
arXiv Machine Learning
Sep 17

Integrated Optimization of Automated Warehouse Operations and Last-Mile Transport for Differentiated On-Demand Delivery

The paper introduces an integrated optimization framework that links automated warehouse operations with last‑mile multi‑modal transport for differentiated on‑demand delivery. It employs a deep reinforcement learning approach—MORM‑AGDQN for warehouse scheduling and MRMH‑HCVRP for external routing—to balance service level, cost, and demand. The results demonstrate significant performance gains, including a 100 % on‑time delivery rate, a 29.3 % reduction in average last‑mile delivery time, a 46.4 % cut in total transportation distance, and a high‑priority service rate exceeding 92 % while maintaining cost‑customer satisfaction balance.

By Xiaozhu Sun, Bilal Farooq
arXiv Machine Learning
Sep 17

The Latent That Never Was: A Forensic Re-run of the CVAE Ablation in Action Chunking Transformers

The paper re‑examines the impact of removing the encoder from Action Chunking Transformers (ACT), a model used for robot manipulation learning. Contrary to the original claim that encoder removal drops success rates from 35% to 2%, the authors find no such dramatic decline in their re‑runs, though minor variations remain uncertain. They attribute the discrepancy to factors like training length and checkpoint selection, and note that the encoder’s latent variable offers little reconstruction benefit on the tested benchmark, while its removal speeds up training.

By Bo Kang
arXiv Computation and Language
Sep 17

CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation

The paper introduces CROP, a method for selective on‑policy distillation that prioritizes token‑level supervision based on task relevance. CROP uses paraphrase‑calibrated counterfactual sensitivity to measure how much each response token depends on the semantic content of the input, constructing validated original‑paraphrase‑counterfactual triplets for each prompt. Experiments in two teacher‑student settings show that CROP outperforms other selectors, improving aggregate performance by 1.92 and 2.96 points.

By Enhan Li, Junhao He, Hongyang Du
arXiv Computation and Language
Sep 17

RT-SEMamba: Real-Time Speech Enhancement Mamba via Progressive Knowledge Distillation

RT-SEMamba is a fully causal speech enhancement model that uses causal time‑frequency Mamba blocks instead of Transformer‑based architectures, allowing efficient long‑form inference with a fixed‑size recurrent state. The authors introduce a progressive knowledge distillation strategy that compresses an 8‑layer teacher into a single‑layer student by jointly distilling spectral outputs and intermediate representations. On the Voicebank‑DEMAND benchmark, the 8‑layer model achieves 3.32 PESQ under a 25 ms latency constraint, while the distilled 1‑layer student improves from 3.06 to 3.18 PESQ, maintains the same steady‑state real‑time factor, and runs 2.64× faster than the teacher.

By Rong Chao, Sung-Feng Huang, Moreno La Quatra, Sabato Marco Siniscalchi, Wen-Huang Cheng, Szu-Wei Fu, Yu Tsao
arXiv AI
Sep 17

Beyond Truncation: Rethinking LLM Decoding as Ensemble Pruning

The paper introduces Mahalanobis-Ensemble Decoding (ME-Decoding), a new framework for Large Language Model decoding that treats candidate token selection as an ensemble pruning problem. It uses a Mahalanobis distance-driven objective to promote semantic diversity while maintaining high probabilities, employing a token similarity matrix built with an adaptive-bandwidth kernel over token embeddings. An efficient greedy algorithm with near-linear complexity and theoretical guarantees makes ME-Decoding a plug‑and‑play module with negligible inference overhead, and experiments show strong performance across reasoning and generation tasks.

By Dunyao Xue, Chengshuo Du, Zhengbo Wang, Wenlin Dai, Cheng Meng
arXiv AI
Sep 17

Higher-order pruning of experts in mixture-of-experts language models

The paper introduces HOPE, a second‑order pruning method for Mixture‑of‑Experts language models that accounts for cooperative interactions between experts. Unlike first‑order methods such as REAP, HOPE derives an objective that provably bounds pruning error and is shown to outperform baselines across three large MoE models, multiple calibration sets, and diverse benchmarks, especially at high pruning rates and on agentic tasks. The results demonstrate that preserving expert interactions allows aggressive compression with minimal performance loss on complex workloads.

By Alex M. Tseng, Prannay Kaul, Luca Zancato, Wei Xia, Stefano Soatto
arXiv Machine Learning
Sep 17

Sparse Bayesian Modeling of EEG Channel Interactions Improves P300 Brain-Computer Interface Performance

The paper introduces a sparse Bayesian time‑varying regression framework that models pairwise EEG channel interactions and performs temporal feature selection for P300 brain‑computer interfaces. Using a relaxed‑thresholded Gaussian process prior, the method achieves a median character‑level accuracy of 96.4% on a public P300 speller dataset and outperforms both statistical and deep learning baselines. It also yields subgroup‑specific gains, notably for participants who abstain from alcohol, and improves median BCI‑Utility by over 10%, reaching peak throughput after only six sequence repetitions.

By Guoxuan Ma, Yuan Zhong, Moyan Li, Yuxiao Nie, Jian Kang
arXiv Computer Vision
Sep 17

Decoder-Agnostic Token Merging for Vision Transformers: A Systematic Study of G2TM

The paper studies Graph-Guided Token Merging (G2TM), a module that reduces token count in Vision Transformers. It evaluates G2TM across multiple segmentation frameworks and decoder types, finding that its performance gains are tied to the encoder rather than the decoder. The authors report consistent reductions in GFLOPs (22‑47%) and throughput improvements (up to 74%) on ADE20K, with optimal hyperparameters depending mainly on backbone pre‑training and target dataset.

By Victor Bercy, Martyna Poreba, Michal Szczepanski, Samia Bouchafa
arXiv Computer Vision
Sep 17

Generalizable Neural Reconstruction of High-Fidelity Surfaces via Sparse Volumetric Representations

The paper introduces SVRecon, a generalizable neural surface reconstruction framework that uses sparse volumetric representations to achieve high-resolution 3D reconstruction. It employs a two-stage architecture: first predicting occupied voxels with an occupancy network, then rendering only within those regions using specialized sparse algorithms. This approach allows reconstruction at resolutions up to 512³ on 32 GB hardware, producing smoother and more precise surfaces, especially in sparse-view scenarios.

By Aoxiang Fan, Corentin Dumery, Nicolas Talabot, Ming Xu, Hieu Le, Pascal Fua
arXiv AI
Sep 17

FairCompressAgent: An Agentic Framework for Fairness-Aware Model Compression for FPGA Deployment

FairCompressAgent (FCA) is an agentic framework that unifies fairness-aware pruning, incremental quantization, and sparse low‑rank factorization for FPGA deployment. A language‑model planner selects compression configurations based on model profiles and measured outcomes, while an execution layer handles compression, fine‑tuning, evaluation, and constraint‑based selection. Experiments on Fitzpatrick‑17k with VGG‑11 show FCA can reduce inference tensor storage by 59.54% under accuracy constraints, improve validation average precision, and lower equalized opportunity, achieving similar results to one‑shot planning with fewer candidate evaluations.

By Yuanbo Guo, Yiyu Shi
arXiv AI
Sep 17

The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction

The paper introduces Edge0, a streaming mixture‑of‑experts (MoE) inference engine that enables a 35‑billion‑parameter MoE model to run on consumer hardware by predicting routing decisions one token ahead. Edge0 uses a per‑layer prerouter to prefetch the necessary experts from SSD, and an unmerged recovery LoRA trained on the student path to recover quality lost to 4‑bit quantization and routing replacement. On a single 24‑GB machine, Edge0 serves the 35B MoE at 20 tokens per second while keeping peak active memory below 3 GiB, achieving performance close to its fp16 teacher across five public benchmarks.

By Yu Lin, Yiming Wang, Runyuan Cai, Hanze Liu, Xiaodong Zeng
arXiv AI
Sep 17

Function Lives Where Variance Doesn't: Task-Weighted Charts of a Language Model's Computation

The paper introduces task‑weighted charts, a method that defines low‑dimensional coordinate systems based on a chosen functional of a language model’s representation. Using these charts, the authors show that next‑token prediction requires 70–90% of the residual stream’s width to maintain perplexity, a width largely unused by variance‑based analyses. The study demonstrates that charts trained under the functional’s metric preserve predictions better than variance‑based or optimal linear compression when only a few dimensions are retained.

By Alexandre Quemy
arXiv AI
Sep 17

REQAP: Resilient Weight Packing and Quantization for Edge DNN Acceleration

The paper introduces REQAP, a reliability‑aware quantized weight packing technique for systolic‑array DNN accelerators. It uses a sensitivity‑driven mixed‑precision quantization to assign layer‑wise bit‑widths, a deterministic register‑level packing strategy for SIMD‑within‑a‑register execution, and selective bit‑level protection that replicates critical MSBs into unused register space. Experiments on AlexNet, VGG‑11, and ResNet‑18 show up to 62% memory reduction, 56% fewer MAC operations, and improved accuracy resilience under fault injection compared to baseline and fully protected models.

By Mahdi Taheri, Samira Nazari, Mubassher Ansari, Ali Azarpeyvand, Mohsen Afsharchi, Maksim Jenihhin, Christian Herglotz
arXiv AI
Sep 17

One Size Does Not Fit All! Dynamic Retriever and Generator Selection for RAG

The paper introduces DRAG, a query‑adaptive framework that jointly selects retriever and generator configurations for Retrieval‑Augmented Generation (RAG) systems. Two variants are presented: DRAG_QPP, a training‑free routing method using Query Performance Prediction and perplexity signals, and DRAG_SFT, a supervised approach that fine‑tunes an LLM to predict configurations. Experiments on three LLM families and four QA benchmarks show that DRAG_QPP matches strong static baselines while cutting inference latency, and DRAG_SFT consistently outperforms both static and training‑free adaptive baselines, demonstrating a better effectiveness‑efficiency trade‑off.

By Neeraj Anand, Payel Santra, Partha Basuchowdhuri, Debasis Ganguly, Sumit Bhatia
arXiv AI
Sep 17

HyQuant: Hybrid-Precision Quantization for LLM Attention

HyQuant introduces a hybrid-precision quantization framework for large language model (LLM) attention modules. It quantizes most attention states to low-bit formats while preserving a small set of vertical‑line tokens and local‑window states in full precision, guided by lightweight attention‑pattern signals. This design achieves near‑lossless accuracy across tasks while improving memory and hardware efficiency.

By Jiatong Ding, Bingxin Xing, Yu Zhang, Dian Ding, Xiaodong Yi, Xianbin Ouyang, Feihu Zhou, Kun Zhang, Zhenyu Guo, Hao Pan, Guangtao Xue, Yiming Zhang