Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,834 stories · RSS feed

arXiv Machine Learning
Sep 22

Glass-Box Deep Learning for FDIA Detection in Nonlinear Automatic Generation Control: A Kolmogorov-Arnold Network Approach

arXiv:2509.05259v2 Announce Type: replace Abstract: Automatic Generation Control (AGC) plays a critical role in maintaining power balance across multi-area power systems. However, its complete relian...

By Ahmad Mohammad Saber, Alok Paranjape, Jehad Jilan, Niranjana Naveen Nambiar, Amr Youssef, Deepa Kundur
arXiv Computer Vision
Sep 22

SparkDiffusion: Mitigating the High-Sparsity Trap --- A Unified Framework for up to $265\times$ Single-GPU Acceleration of Visual Generation

arXiv:2609.23153v1 Announce Type: new Abstract: Video diffusion transformers are expensive because attention dominates long spatiotemporal token sequences. We identify the \emph{high-sparsity trap}:...

By Yuxi Liu, Haoyu Li, Zekun Zhang, Tengxu Sun, Yixiang Cai, Jiayong Li, Yifei Xia, Tianle Liu, Baole Ai, Ang Wang, Jiamang Wang, Lin Qu, Kai Zhang, Kun Yuan, Bin Cui
arXiv Machine Learning
Sep 22

Helix-FNO: Spectral-Domain Operator Learning Coupled with a High-Fidelity Mechanistic Model for Fast Surrogate Simulation

Helix‑FNO is a teacher‑student framework that couples a 32‑state mechanistic model with a Fourier neural operator to learn the full solution operator for full‑scale treatment processes. The teacher generates a high‑fidelity dataset via Latin‑hypercube sampling and active learning, while the student learns in the spectral domain, enabling generalisation across varying influent profiles, controls, and plant layouts. The resulting operator achieves millisecond inference, three orders of magnitude faster than the mechanistic teacher, and is evaluated against physics‑informed and data‑driven surrogates on accuracy, dataset efficiency, and latency, positioning it on a speed‑accuracy Pareto front.

By Jiabao Zhao, Chuwei Wang, Jinxi Yang
arXiv Machine Learning
Sep 22

Tool-Augmented On-Policy Distillation for LLM Domain Adaptation in Sequence-Based Omics Tasks

The paper introduces OmicsBench, a new reasoning benchmark for multi‑omics sequences that includes 1,160 expert‑validated questions across DNA regulation, RNA processing, and protein function tasks, requiring traceable evidence chains. Evaluation of 17 large language models shows that scientific LLMs, while more accurate in classification, often lack valid evidence, suggesting shortcut learning. To address this, the authors propose tool‑augmented on‑policy distillation (TA‑OPD), a post‑training method that improves both evidence grounding and predictive performance across five Qwen3.5 models of varying sizes.

By Jie Ying, Zhefan Wang, Zihong Chen, Zhengqing Li, Jinzhe Li, Gang Li, Jian Liu, Fang Hu, Tao Luo, Zhonghang Yuan, Wanli Ouyang, Stan Z. Li, Fan Yang, Nanqing Dong