Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,633 stories · RSS feed

arXiv Machine Learning
2d ago

Surprisingly High Redundancy in Electronic Structure Data Across Materials Explained by Low Intrinsic Dimensionality

arXiv:2507.09001v4 Announce Type: replace-cross Abstract: Machine learning (ML) models for electronic structure typically rely on large datasets generated by computationally expensive Kohn-Sham densi...

By Sazzad Hossain, Ponkrshnan Thiagarajan, Shashank Pathrudkar, Stephanie Taylor, Abhijeet S. Gangan, Amartya S. Banerjee, Susanta Ghosh
arXiv Computer Vision
2d ago

PAGER: Partial-to-global Alignment via Geometric and Relational Distillation

arXiv:2610.01589v1 Announce Type: new Abstract: Pretrained 3D encoders are typically developed on globally reconstructed scenes expressed in a consistent world coordinate frame, whereas embodied syst...

By Akira-Miranda Adeyomi Adeniran-Lowe, Binod Singh, Lars Arnold Dethlefsen, Lazaros Nalpantidis, Theodora Kontogianni
arXiv AI
2d ago

HHR: Hierarchical Hash Retrieval for Efficient LLM Generation

The paper introduces Hierarchical Hash Retrieval (HHR), a coarse‑to‑fine framework designed to improve hash‑based retrieval for large language models. HHR combines Geometry‑Aware Key Routing (GKR) to redistribute feature magnitudes and prune low‑logit keys, with Learned Hash Projection (LHP) to align Hamming distance with true query‑key relevance for fine‑grained retrieval. Experiments on diverse LLMs and benchmarks show that HHR outperforms existing methods, boosting LongBench scores by 1.10 points and achieving up to 3.30× decoding speedup at 128K context length for Llama‑3.1‑8B‑Instruct.

By Lianjun Liu, Tiantian Zheng, You Huang, Weiqi Yan, Mingte Qiu, Huazhong Liu, Xiaofeng Zhu, Yunshan Zhong
arXiv AI
2d ago

MOMAT: Mixture of Multiple Atlases for Low-Power Jailbreak Defense of Quantized LLMs

MOMAT (Mixture of Multiple Atlases) is a hardware‑enhanced safety framework designed to defend quantized large language models (qLLMs) on edge devices against jailbreak attacks. It uses a collection of semantic atlases—each containing harmful or benign sample clusters and policy templates—to perform domain‑localized Retrieval‑Augmented Generation. A lightweight Mixture of Experts detector evaluates top‑k similarity features retrieved by a Compute‑in‑Memory (CiM) accelerated engine, achieving a 4.69 × 10⁶‑fold speedup and a 2.5 × 10⁵‑fold energy reduction compared to DRAM‑based baselines while matching state‑of‑the‑art defense performance.

By Boyang Li, Bingyu Shen, Weihao Hong, Zhiyuan Jiang, Xinlei Guan, Yan Ma, Miles Q. Li, Yi Sheng, Ruiyang Qin
arXiv Machine Learning
2d ago

Decoupled and Distilled: Task-Adaptive LoRA-Teachers with Ensemble Knowledge Transfer for Few-Shot Class-Incremental Learning

The paper introduces TALON, a Task‑Adaptive LoRA‑Teacher framework for Few‑Shot Class‑Incremental Learning. TALON assigns a dedicated LoRA‑Teacher to each incremental task, then distills the frozen teachers into a single LoRA‑Student via Ensemble Knowledge Transfer, using a semantic‑guided weighting scheme to reduce forgetting and overfitting. Experiments on four FSCIL benchmarks show that TALON matches or surpasses state‑of‑the‑art accuracy while using up to 33× fewer deployment parameters and cutting inference time by 41.7%.

By Hongwei Zhao (School of Computer Science,Engineering, Beihang University), Rui Liu (School of Computer Science,Engineering, Beihang University), Yansong Liu (School of Computer Science,Engineering, Beihang University), Zhiyuan Zou (School of Computer Science,Engineering, Beihang University), Yong Chen (School of Computer Science, Beijing University of Posts,Telecommunications)
arXiv Machine Learning
2d ago

RATIO: Reasoning Analysis and Token-level Inference Optimization for Quantized Reasoning Models

The paper introduces RATIO, a framework for improving quantized reasoning models by identifying overthinking tokens and applying token-specific penalties. It uses Quantization-aware Reasoning Behavior Analysis to detect problematic tokens and Token-Specific Penalty Determination to assign penalties without extra training. Experiments show RATIO outperforms existing methods, boosting accuracy by up to 9.8 points and shortening chain-of-thought length by up to 51.3%.

By Chengzhu Bao, Xianglong Yan, Tianao Zhang, Jiaqi Chen, Shaoqiu Zhang, Yulun Zhang
arXiv AI
2d ago

Fault-Tolerant Budget Conservation in Distributed Multi-Agent Delegation

The paper introduces a fault‑tolerant budget conservation framework for distributed multi‑agent delegation, where budgets are represented as exclusive escrow credits that traverse a delegation DAG. It details how each branch converts credit into a reservation tied to lineage, epoch, and idempotency, persists a signed dispatch permit, and ensures that uncertain effects remain charged until settlement or retirement. The authors prove properties such as ownership partition, ledger conservation, and at‑most‑once settlement, and validate the mechanism through TLA+ checks, a JavaScript explorer, and crash‑injected SQLite experiments.

By Genliang Zhu, Chu Wang
arXiv AI
2d ago

ProtoFlow: Prototype-Guided Flow Matching for Multivariate Time Series Forecasting

ProtoFlow is a new multivariate time series forecasting framework that combines vector‑quantized autoencoding with prototype‑guided flow matching. It maps sequences into a discrete latent space, constructs a structured prior from the learned VQ codebook, and trains a DiT‑based rectified flow to transport samples from this prior to future latent representations conditioned on past observations. By replacing generic Gaussian noise with a learned prototype prior, ProtoFlow eliminates autoregressive rollout mismatch and achieves faster training convergence while delivering superior forecasting performance on benchmark datasets.

By Shibo Feng, Wanjin Feng, Yang Qiu, Deheng Ye, Peilin Zhao, Chunyan Miao
arXiv AI
2d ago

Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing

The paper presents a billing‑aware neural text‑to‑speech system for serverless CPUs, focusing on minimizing CPU‑seconds and GB‑seconds rather than just throughput or latency. By using request‑sized concurrent inference and a reclaimable instance lifecycle, the system limits per‑request CPU parallelism and releases idle memory while keeping the server process alive. On the Kokoro‑82M benchmark, it achieves 2.71 audio‑seconds per CPU‑second versus 0.89 with ONNX Runtime, cuts cost per audio‑hour from $0.0631 to $0.0153, and reduces idle billed memory from 8.7 GB to 1.33 GB, with faster restoration times.

By Pakorn Nathong, Kunat Pipatanakul
arXiv AI
2d ago

Per-Node Activation Function Evolution in Indirectly Encoded Substrates: Solvability, Limits, and Emergent Diversity

The paper demonstrates that using a single activation function across all nodes in artificial neural networks imposes hard limits on evolutionary search, particularly for sparse evolved substrates. By evolving per-node activation functions from an 18-function palette, the authors show that oscillatory functions can solve parity problems at all tested scales, while monotonic functions fail beyond the simplest case. The study reveals that the choice of activation functions, beyond topology and weights, critically influences what evolutionary search can achieve, and that heterogeneous assignments discovered via indirect encoding are unlikely to be selected manually.

By Romain Claret, Michael O'Neill, Paul Cotofrei, Kilian Stoffel
arXiv AI
2d ago

Four Ways to Grow a Classifier and Why One of Them Cannot Learn

The paper investigates four ways to grow a classifier—adding a tree level, a hidden unit, a leaf split, and a statistically significant split—under a fixed protocol for tree‑structured and constructive models. It shows that the most natural method of deepening a soft decision tree by duplicating a leaf’s class distribution leaves the gradient of new gates identically zero, preventing learning, and proposes a small random perturbation as a fix. The other three growth decisions each provide a distinct benefit: fitting a new hidden unit to residual error yields a smaller network, splitting the leaf with the largest expected error adds sparsity, and requiring statistical significance before splitting adds no value and reduces accuracy.

By Cagri Temel
arXiv AI
2d ago

DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies

DriftOPD is a teacher‑free, rollout‑free framework that performs sequence‑level on‑policy distillation of continuous Vision‑Language‑Action (VLA) action experts. It decomposes the sequence‑level reverse‑KL divergence into a chunk‑level reverse‑KL term and a future‑potential term, optimizing them with a one‑step drifting objective and a Q‑function critic learned from offline demonstrations. Experiments on multiple VLA architectures in simulation and real‑world manipulation show that DriftOPD outperforms existing one‑step distillation baselines while matching the task success of multi‑step teacher policies.

By Youngjun Jun, Kyumin Choi, Youngmin Kim, Seonghyun Jin, Sunwoo Park, Jangho Park, Jong Chul Ye