Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,914 stories · RSS feed

arXiv AI
Sep 16

Calibrate, Then Route: A Measured Study of Learned Request Routing for Disaggregated LLM Serving

The paper evaluates a learned request‑routing policy for disaggregated large‑language‑model serving, where compute‑heavy prefill and memory‑heavy decode stages run on separate GPU pools. Using a discrete‑event simulator and real NVIDIA A40 GPUs, the calibrated router—leveraging prompt length, predicted output length, KV‑cache pressure, and SLO class—outperforms round‑robin, least‑loaded, and length‑based heuristics, achieving the highest mean goodput (0.864) and lowest variance across three mixed, bursty arrival traces. Hardware calibration proves critical, providing a 4.5‑point goodput boost and roughly 40 % of the tail‑latency advantage, and the learned router can match round‑robin performance with one fewer GPU in certain scenarios.

By Srikanta Datta Tumkur, Jay Iyer, Mehar Simhadri, Sai Pavan Kumar, Sai Kapil Kumar, Ramesh Nampelly
arXiv Machine Learning
Sep 16

The Latent That Never Was: A Forensic Re-run of the CVAE Ablation in Action Chunking Transformer

The paper re‑examines the impact of removing the encoder from Action Chunking Transformers (ACT), a model used for robot manipulation learning. Contrary to the original claim that encoder removal drops success rates from 35% to 2%, the authors find no such dramatic effect in their re‑runs, though minor variations remain uncertain. They attribute discrepancies to training length and checkpoint selection, and note that the encoder’s latent variable offers little reconstruction benefit on the tested benchmark, while its removal speeds up training.

By Bo Kang
arXiv Machine Learning
Sep 16

MedPCFM-TED: One-Step Point Cloud Flow Matching for Implant Generation via Teacher-Guided Endpoint Distillation

The paper introduces MedPCFM‑TED, a one‑step distillation framework that uses teacher‑guided endpoint supervision and geometric matching losses to generate cranial implants from point clouds. It outperforms existing one‑step methods on the SkullBreak benchmark, remains competitive on SkullFix, and achieves a generation time of about 0.04 s per sample. The approach demonstrates that rapid, high‑quality implant generation is possible without multiple neural evaluations during inference.

By Kamil Kwarciak, Marek Wodzinski
arXiv Machine Learning
Sep 16

LCAP: Population-Informed Latent Chip Adaptation from Few Output Probes for Photonic Neural Networks

The paper introduces LCAP, a method for adapting photonic neural networks to real hardware by learning a shared correction from a population of chips and then personalizing each chip using only 32 fixed output probes. LCAP decomposes adaptation into a transferable population correction and a probe‑inferred latent personalization, allowing feed‑forward calibration without device‑specific optimization. Experiments on a simulated three‑layer 64‑mode MZI network show accuracy improvements from 80.4% to 93.4% and significant gains on unseen chips.

By Tianyu Gao, Guantian Zheng
arXiv Computation and Language
Sep 16

What Breaks Under Pruning in Smart Homes, and When? Evaluating LLM Degradation Across Architectures and Task Complexity

The study investigates how pruning affects large language models (LLMs) used for smart‑home tool calling. Researchers examined four LLMs—dense Transformer, dense hybrid, and mixture‑of‑experts (MoE) architectures—using depth, width, hybrid, and expert pruning, followed by supervised fine‑tuning. They evaluated over 19,500 instances from three smart‑home datasets, analyzing not only overall accuracy but also degradation in action components (operation, device, argument, value) and task complexity, finding that dense models suffer sharp performance drops after a narrow safe pruning range, while MoE models tolerate more pruning; aggressive pruning also leads to over‑refusal and loss of grounded specificity.

By Congjing Zhang, Vashishtha Patil, Henning Lange, Usman Aleem
arXiv Computer Vision
Sep 16

Accelerated Decoding of Centroid Positional Encoding for Instance Segmentation

arXiv:2609.16874v1 Announce Type: new Abstract: Beyond model inference, the decoding stage, which converts raw network outputs into task-level representations, constitutes a significant portion of th...

By Carmelo Scribano, Filippo Muzzini, Nedyalko Prisadnikov, Mohammad Mahdi, Yuqian Fu, Giorgia Franchini, Danda Pani Paudel, Marko Bertogna, Luc Van Gool
arXiv Machine Learning
Sep 16

LLM Inference in a Flash!

The paper "LLM Inference in a Flash!" proposes an integer‑only quantization scheme and a dictionary‑based KV cache compression technique to enable large language model inference on compute‑in‑flash (CIF) devices. By eliminating floating‑point operations and reducing KV cache traffic through sparse dictionary coding, the authors achieve minimal accuracy loss while cutting dynamic KV cache traffic by 15× on Llama‑3.1‑8B and Qwen‑2.5‑7B models.

By Sebastian Zhao, Minseo Kim, Coleman Hooper, Luca Manolache, Michael W. Mahoney, Yakun Sophia Shao, Kurt Keutzer, Amir Gholami