arXiv Machine Learning By Zelin Li, Caiwen Ding

Inference-Native Zeroth-Order Optimization

Read the original on arXiv Machine Learning →

The paper introduces Inference‑Native Zeroth‑Order (ZO) optimization, which redefines ZO as a query‑based process that can be executed directly by inference runtimes. By exposing ZO’s query semantics and using abstractions such as ProbePlan, factorized side states, and persistent subspace reuse, the method reduces state‑management cost and DRAM traffic dramatically. Experiments on large models (OPT‑13B, Qwen3‑8B) show that inference‑native steps are nearly identical to matched‑query controls while achieving significant memory savings and efficient batching.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 11

Optimizing AI Inference Across the Deployment Stack

The paper argues that AI deployment performance depends on interactions among compression, compiler transformations, and serving policies rather than just model architecture. It introduces a three‑layer taxonomy—model‑level techniques, compiler transformations, and system policies—and frames deployment as a constrained multi‑objective optimization problem over accuracy, latency, throughput, memory footprint, and energy. The authors propose an evidence protocol for comparable benchmarking and synthesize data from edge and data‑center platforms to show that cross‑layer interactions drive deployment outcomes, concluding with a constraint‑aware selection procedure and open research problems.

By Tejinder Singh, John Pflueger, Jeebak Mitra, Robert Lincourt, Mitchell Markow, Bhavesh A. Patel
arXiv Machine Learning
Sep 14

Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference

Dynamic Expert Quantization (DynaExq) is a runtime-aware mixed-precision serving system designed for single‑GPU Mixture‑of‑Experts (MoE) inference under a hard high‑bandwidth memory (HBM) envelope. It treats the problem as an online, budget‑constrained precision allocation task, keeping the most frequently used experts at higher precision while relegating the rest to low‑precision fallbacks. By estimating expert hotness from router traces and asynchronously promoting or demoting experts, DynaExq maintains a fully materialized expert set during the forward pass, improving accuracy and throughput compared to static post‑training quantization and offloading/prefetch baselines. whyItMatters":"DynaExq enables efficient deployment of large MoE models on memory‑limited GPUs by dynamically allocating precision based on runtime expert usage, thereby reducing memory footprint and latency while boosting accuracy and throughput."

By Kexin Chu, Dawei Xiang, Zixu Shen, Yiwei Yang, Zecheng Liu, Wei Zhang
arXiv AI
Aug 28

Compositional Online Learning for Semantic Data Processing Systems

The paper introduces compositional online learning for semantic data processing systems, addressing the high cost and latency of large language model (LLM) calls. It proposes a framework that combines lightweight online-learning components—such as memoization, per-call filter-ordering, and per-batch cascade-routing—within the LLM call boundary, allowing each component to make real-time decisions and update its models without exceeding the LLM round-trip time. A production case study in Cortex AISQL demonstrates that these components can reduce the per-row LLM cost by up to 8× compared to a baseline workload.

By Pawe\l{} Liskowski, Fuheng Zhao, Benjamin Han, Anupam Datta, Dimitris Tsirogiannis