Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

6,032 stories · RSS feed

arXiv Machine Learning
Sep 14

Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference

Dynamic Expert Quantization (DynaExq) is a runtime-aware mixed-precision serving system designed for single‑GPU Mixture‑of‑Experts (MoE) inference under a hard high‑bandwidth memory (HBM) envelope. It treats the problem as an online, budget‑constrained precision allocation task, keeping the most frequently used experts at higher precision while relegating the rest to low‑precision fallbacks. By estimating expert hotness from router traces and asynchronously promoting or demoting experts, DynaExq maintains a fully materialized expert set during the forward pass, improving accuracy and throughput compared to static post‑training quantization and offloading/prefetch baselines. whyItMatters":"DynaExq enables efficient deployment of large MoE models on memory‑limited GPUs by dynamically allocating precision based on runtime expert usage, thereby reducing memory footprint and latency while boosting accuracy and throughput."

By Kexin Chu, Dawei Xiang, Zixu Shen, Yiwei Yang, Zecheng Liu, Wei Zhang
arXiv Computer Vision
Sep 14

Semantically Aligned Gradient-Driven Context-Preserving Image Editing

Semantically Aligned Gradient-Driven Context-Preserving Image Editing (IABEdit) is a model‑agnostic framework that embeds differentiable semantic verification into the training of generative image editors. By using a frozen vision‑language model to extract spatially‑aware descriptors from ground‑truth edits and a trainable aligner to reproduce them from generated outputs, the residual becomes a gradient that teaches the generator both what to edit and where, without adding inference‑time VLM cost. IABEdit is compatible with various backbones (e.g., U‑Net in Stable Diffusion and MMDiT in FLUX) and improves structural fidelity on MagicBrush, achieves state‑of‑the‑art instruction adherence on RealEdit and EMU Edit, and outperforms the proprietary Gemini agent on the D‑LORD surveillance benchmark under heavy occlusion. "whyItMatters":"IABEdit demonstrates that incorporating semantic verification during training can produce more accurate, well‑localized edits and outperform existing methods even in challenging surveillance scenarios, as shown by its superior metrics and human/GPT‑4o evaluations."

By Chiranjeev Chiranjeev, Muskan Dosi, Mayank Vatsa, Richa Singh
arXiv Machine Learning
Sep 14

FEAT: A Linear-Complexity Foundation Model for Extremely Large Structured Data

FEAT is a foundation model designed for extremely large structured data that replaces quadratic self‑attention with a linear‑complexity dual‑axis encoding architecture. It combines an adaptive‑fusion bidirectional state‑space model with convolutional gated linear attention to achieve permutation‑invariant representation learning in O(N) time. Experiments on 12 real‑world database benchmarks show that FEAT outperforms existing structured data foundation models on zero‑shot tasks and can be up to 50× faster in inference latency.

By Zhenghang Song, Tang Qian, Lu Chen, Yushuai Li, Zhengke Hu, Bingbing Fang, Yumeng Song, Junbo Zhao, Sheng Zhang, Tianyi Li
arXiv Computer Vision
Sep 14

Uni-HOI:A Unified framework for Learning the Joint distribution of Text and Human-Object Interaction

Uni-HOI is a unified framework that learns the joint distribution among text, human motion, and object motion for 4D human‑object interaction (HOI). It uses large language models and two motion‑specific VQ‑VAEs to convert heterogeneous motion data into token sequences, enabling seamless integration of all three modalities. A two‑stage training strategy first captures correlations on a large‑scale HOI dataset and then fine‑tunes for specific tasks, achieving strong performance on text‑driven HOI generation, object‑motion‑driven human motion generation, and human‑motion‑driven object motion prediction.

By Mengfei Zhang, Jinlu Zhang, Zhigang Tu
arXiv Machine Learning
Sep 14

Theoretical Guarantees for One-Shot Magnitude Pruning and Compute-Adaptive Early Exit

The paper investigates how to reduce computation in neural networks by combining one‑shot magnitude pruning in a static setting with early exit in an adaptive setting. In a simplified single‑neuron model it proves a concentration theorem for pruning and introduces a conditional perceptron whose excess error decreases as a power of the compute gap, with the exponent increasing as partial and full computations align. The authors extend these results to deep networks, showing how pruning distortions accumulate with depth and deriving a compute‑accuracy trade‑off for frozen‑backbone early exit under a Gaussian process framework, with numerical simulations supporting the theoretical scaling laws.

By Erdem Koyuncu
arXiv Machine Learning
Sep 14

Representation Before Training: A Practical Benchmark for Generative Medical Event Model Tokenization

The study benchmarks tokenization choices for generative medical event models, evaluating quantization granularity, reference-range anchoring, code–value fusion, numeric and temporal encodings, and native versus harmonized event representations. Using Llama and Qwen architectures, 156 models were trained and assessed on early hospitalization data, showing that fusing codes with value deciles and using event-order or admission-relative RoPE embeddings improved predictive performance. The Common Longitudinal Intensive Care Unit Data Format (CLIF) reduced token count by 30.8% while enhancing outcomes in most families.

By Inhyeok Lee, Luke Solo, Michael C. Burkhart, Bashar Ramadan, Sahil Sethi, Sarah Jabbour, William F. Parker, Brett K. Beaulieu-Jones
arXiv Machine Learning
Sep 14

Dual-guided Hierarchical Edge Localization for Large-scale Optimal Transport Across Dimensions

The paper introduces HELLO, a hierarchical solver for large‑scale discrete optimal transport that reduces the problem to edge localization guided by dual potentials. HELLO uses a coarse‑to‑fine initialization across a recursive subsampling hierarchy and a refinement step that inserts the largest dual violators until a KKT residual tolerance is met, achieving linear memory usage. Experiments show that HELLO outperforms strong baselines by an order of magnitude in runtime while attaining lower transport objectives, and it scales to over a million samples in high‑dimensional settings, supporting various OT variants.

By Wenzhou Xia, Qiaoqiao Ding, Jingwei Liang, Xiaoqun Zhang
arXiv Machine Learning
Sep 14

Block-Norm Geometries for Online Mirror Descent with Sparse Losses

The paper investigates how the choice of geometry in online mirror descent affects performance, particularly when loss gradients are sparse. It introduces randomized block‑norm mirror maps that interpolate between Euclidean and entropic geometries, achieving polynomial‑in‑dimension regret improvements over standard methods for various convex sets. The authors also demonstrate that naive alternation between mirror maps can lead to linear regret and propose a Hedge‑based meta‑algorithm that competes with the best mirror map in a finite portfolio, achieving near‑optimal regret for random block geometries.

By Swati Gupta, Jai Moondra, Mohit Singh
arXiv Computation and Language
Sep 14

Parameter-Efficient Retrievers for Polish and European Languages

The paper introduces a three‑stage training pipeline that builds compact, efficient dense retrievers without requiring ground‑truth relevance labels. Using cross‑lingual alignment, relational knowledge distillation, and contrastive fine‑tuning, the authors develop PolDense (six Polish models ranging from 17 M to 1 B parameters) and EuroDense (a 435 M‑parameter model covering nine European languages). Extensive evaluation on 41 Polish and 150 multilingual tasks shows that PolDense‑1B outperforms larger retrievers up to 9 B parameters, while EuroDense leads in task‑averaged and language‑averaged performance among models below 1 B parameters.

By S{\l}awomir Dadas, Rafa{\l} Po\'swiata, Ma{\l}gorzata Gr\k{e}bowiec, Micha{\l} Pere{\l}kiewicz
arXiv Machine Learning
Sep 14

Quality-Constrained Routing over a Fixed Pool of Quantized Mixture-of-Experts Instances

The paper proposes a method for routing requests to a fixed pool of quantized Mixture-of-Experts (MoE) instances, aiming to maximize throughput while respecting a quality‑degradation budget. It introduces Fragility‑Weighted Perplexity (FWP) as a request‑specific risk metric derived from prompt tokens, and uses a window‑level linear program to compute a reduced‑reward score that aligns with the LP optimum. Experiments on Qwen prompts show that FWP‑based allocation improves throughput by 2.5% over request‑agnostic mixing and static configurations.

By Zhenghong Huang, Hongfan Wu, Jiheng Zhang
arXiv Computation and Language
Sep 14

SynthSentry: Detecting Synthetic Data Contamination in Language Model Training Data

The paper introduces SynthSentry, a model‑agnostic method for detecting synthetic data contamination in language‑model training corpora. It computes a distributional divergence score based on lexical diversity collapse, n‑gram tail truncation, and perplexity variance across reference models, requiring no access to the generating model or synthetic labels. Experiments on English corpora contaminated by small open‑weight generators and an instruction‑tuned model show that SynthSentry ranks contamination severity accurately, maintains low false‑positive rates after calibration, and does not degrade downstream fine‑tuning performance at the tested scale.

By Praveen Kumar Myakala, Ravichandra Namburi, Sowmya Keragodu Jayaramu, Sooraj George Thomas
arXiv Machine Learning
Sep 14

Bridging Vision Foundation Model Priors with CLIP for Spatial-aware Few-shot Anomaly Detection in Medical Images

The paper introduces Spatial‑FAD, a few‑shot medical anomaly detection framework that fuses Vision‑Language Model (CLIP) semantics with spatial priors from Vision Foundation Models (DINO). A VFM‑enhanced adapter injects structural affinity into CLIP features, while a sliding‑window aggregation produces high‑resolution embeddings for finer lesion localization. Prototype‑enhanced support memory further improves efficiency and performance, yielding significant gains on Liver CT, Retinal OCT, and Brain MRI datasets, notably an 11.4% Dice improvement in 4‑shot scenarios.

By Juzheng Miao, Yuchen Yuan, Cheng Chen, Pheng-Ann Heng
arXiv Machine Learning
Sep 14

The Battery Price of edge AI: A study of the Environmental Impact of LLM Inference on Mobile Devices

The paper investigates the environmental impact of running large language models (LLMs) on mobile devices. It evaluates 18 different LLM configurations on two smartphones and a server, measuring energy per token, latency, accuracy, and battery-cycle consumption. Findings reveal that on-device inference is about three times less energy‑efficient than batched server inference, that energy consumption varies non‑monotonically with quantization bit‑width, and that most models are not on the Pareto front of accuracy and energy efficiency. The study concludes that local AI is not inherently more sustainable than cloud inference, with the majority of environmental impact stemming from device embodied carbon.

By \'Edouard Gu\'egain, Tristan Coignion
arXiv Machine Learning
Sep 14

Efficient AI Model Deployment Using Quantization Analysis Tool

The paper introduces the Quantization Analysis Tool, a system built on the ONNX framework that streamlines quantization workflows for deep learning models. It offers layer‑wise sensitivity analysis, visualizations of weight and activation distributions, and guidance for selecting precision levels to balance model size, latency, and accuracy. Experiments on various neural network architectures show that the tool improves quantized accuracy and overall deployment efficiency.

By Dwith Chenna, Kanishka Macherla
arXiv Machine Learning
Sep 14

On-Device Language Models for Privacy-Preserving Stress Prediction: A Multimodal Evaluation on Mobile Health

The paper investigates on-device language models (ODLMs) for predicting stress in a mobile health context, focusing on privacy-preserving, cloud-independent inference. Using zero‑shot prompting, the authors evaluate ODLMs across multimodal data—objective sensor features and subjective self‑reports—measuring predictive accuracy, latency, and throughput. Results indicate that sensor features slightly outperform self‑reports, and that lightweight sub‑2B models deliver low latency with predictable resource usage, underscoring both the potential and practical limits of ODLMs for mobile mental health.

By Ibukunoluwa Soyebo, Alyssa Donawa, Rodrigo Aguilar Barrios, Brice Patchou, Corey E. Baker
arXiv Machine Learning
Sep 14

Fixed State, Long Reach: What a Constant-Size Cache Buys Block Diffusion at Scale

The paper investigates how block‑diffusion language models can use a constant‑size cache to enable efficient parallel decoding. By employing sequence mixers that summarize completed blocks into a reusable state and a block‑causal training objective, the authors pretrain three 3B block‑diffusion denoisers (attention, Mamba, and hybrid) on 300 B tokens. The resulting state‑space cache remains O(1) in memory and latency regardless of context length, yielding significant speed‑up and memory savings compared to traditional attention‑based caches, especially at very long sequences.

By Vaibhav Singh, Pierre-Andr\'e No\"el, Torsten Scholak, Eugene Belilovsky, Oleksiy Ostapenko
arXiv Machine Learning
Sep 14

A unified self-supervised framework for single-frame Fresnel CDI and overlapped ptychography

The paper introduces a self‑supervised neural network that unifies single‑frame Fresnel coherent diffraction imaging (CDI) and overlapped ptychography. By using a fixed, pre‑estimated probe and optimizing with a Poisson negative log‑likelihood objective, the method reconstructs object patches from either a single diffraction frame or multiple overlapping measurements, achieving high SSIM scores and a ten‑fold improvement in photon‑dose efficiency. Demonstrations on synthetic patterns and real datasets from APS and LCLS show robust, high‑throughput reconstructions, with a 36× speedup over iterative solvers for a 10,304‑frame workload.

By Oliver Hoidn, Steven Henke, Albert Vong, Aashwin Mishra, Apurva Mehta, Matthew Seaberg
arXiv Computer Vision
Sep 14

Learning Sparse Latent Predictive Foundation Model for Multimodal Neuroimaging

The paper introduces Neuro‑JEPA, a sparse multimodal foundation model that learns unified representations of brain MRI across T1w, T2w, and FLAIR sequences using a latent predictive objective and a Mixture‑of‑Experts architecture. It was pretrained on over 1.5 million scans from 428,647 studies and systematically evaluates architectural, masking, objective, and sparsity choices for robust multimodal representation learning. Across 47 tasks from three health systems and 12 public datasets, Neuro‑JEPA consistently outperforms a simple CNN baseline, demonstrating its effectiveness for both clinical and research applications.

By Haoxu Huang, Long Chen, Jingyun Chen, Jinu Hyun, James Ryan Loftus, Kara Melmed, Daniel Orringer, Jennifer Frontera, Seena Dehkharghani, Arjun Masurkar, Narges Razavian