Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

6,032 stories · RSS feed

arXiv Machine Learning
Sep 16

Skeletal Prototypes on Iterative Nerve Expansions

The paper introduces Skeletal Prototypes on Iterative Nerve Expansions (SPINE), a prototype reduction method that represents each class as an embedded 1‑complex rather than a finite set of points. SPINE constructs its initial edge set from a class‑conditional Mapper graph, then refines vertex positions under a classification objective, allowing observations to be assigned to the nearest complex. Evaluated on seventeen benchmark datasets with stratified 10‑fold cross‑validation, SPINE achieves the highest mean accuracy and best average rank among seven competing methods, showing significant improvements over five of them and competitive performance across varying prototype budgets.

By Jordan Eckert, Henry Schenck
arXiv Machine Learning
Sep 16

LLM Inference in a Flash!

The paper "LLM Inference in a Flash!" proposes an integer‑only quantization scheme and a dictionary‑based KV cache compression technique to enable large language model inference on compute‑in‑flash (CIF) devices. By eliminating floating‑point operations and reducing KV cache traffic through sparse dictionary coding, the authors achieve minimal accuracy loss while cutting dynamic KV cache traffic by 15× on Llama‑3.1‑8B and Qwen‑2.5‑7B models.

By Sebastian Zhao, Minseo Kim, Coleman Hooper, Luca Manolache, Michael W. Mahoney, Yakun Sophia Shao, Kurt Keutzer, Amir Gholami
arXiv Computer Vision
Sep 16

Accelerated Decoding of Centroid Positional Encoding for Instance Segmentation

arXiv:2609.16874v1 Announce Type: new Abstract: Beyond model inference, the decoding stage, which converts raw network outputs into task-level representations, constitutes a significant portion of th...

By Carmelo Scribano, Filippo Muzzini, Nedyalko Prisadnikov, Mohammad Mahdi, Yuqian Fu, Giorgia Franchini, Danda Pani Paudel, Marko Bertogna, Luc Van Gool
arXiv Computation and Language
Sep 16

What Breaks Under Pruning in Smart Homes, and When? Evaluating LLM Degradation Across Architectures and Task Complexity

The study investigates how pruning affects large language models (LLMs) used for smart‑home tool calling. Researchers examined four LLMs—dense Transformer, dense hybrid, and mixture‑of‑experts (MoE) architectures—using depth, width, hybrid, and expert pruning, followed by supervised fine‑tuning. They evaluated over 19,500 instances from three smart‑home datasets, analyzing not only overall accuracy but also degradation in action components (operation, device, argument, value) and task complexity, finding that dense models suffer sharp performance drops after a narrow safe pruning range, while MoE models tolerate more pruning; aggressive pruning also leads to over‑refusal and loss of grounded specificity.

By Congjing Zhang, Vashishtha Patil, Henning Lange, Usman Aleem
arXiv Machine Learning
Sep 16

EBL: Efficient Broad Learning for Distributed Adaptive Harmonic Analysis

The paper introduces EBL, an Efficient Broad Learning framework designed for distributed adaptive harmonic estimation in power grids affected by electric vehicle charging. It leverages a quantised FPGA implementation to provide high‑accuracy, half‑cycle input harmonic predictions with ultra‑low latency, outperforming existing FPGA methods by 17.4×. The online transfer learning component enables rapid adaptation across multiple charging scenarios, while bespoke quantisation and sparsity reduce resource usage to just 5.9% of the LUTs on a Zynq Ultrascale+ FPGA, compared to 82% of the state‑of‑the‑art accelerator.

By Changhong Li, Georgios Floros, Biswajit Basu, Shreejith Shanker