Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

6,032 stories · RSS feed

arXiv Machine Learning
Sep 11

Positional task conditioning for scalable defect detection across product families in large product catalogs

The paper presents a method called Positional Task Conditioning (PTC) to improve defect detection in large product catalogs. By breaking detection into focused sub‑tasks and reinforcing task identity at prompt boundaries, PTC reduces context length and isolates error types, boosting F1 scores from 52% to 87%. The approach outperforms rationale‑based distillation across multiple models, achieving near‑state‑of‑the‑art performance at up to 98% lower cost and is deployed in several countries handling over 10 million product families.

By Soham Satyadharma, Gabriel Roccabruna, Suleiman A. Khan
arXiv Machine Learning
Sep 11

SparseDitto: An Agentic Sparse Compilation Framework through Architecture-Aware Synthesis on GPUs

SparseDitto is an agentic sparse compilation framework that jointly synthesizes representation, execution schedule, and hardware mapping for sparse matrix computations on GPUs. It uses structural analysis, a learned template-ranking prior, and LLM-guided lowering to generate CUDA code, with target-GPU profiling refining the plan. The framework supports multiple operators such as SpMV, SpMM, and SpGEMM, adapts to different hardware, and achieves significant speedups over cuSPARSE, including up to 146.61× on certain matrices and 3.39× acceleration for full-batch GCN training.

By Shiyang Li, Guangyan Sun, Jinwei Tang, Yanzhi Wang, Mingyi Hong, Caiwen Ding
arXiv Machine Learning
Sep 11

Negative Self-Distillation: Learning to Reason by Avoiding Flaws

Negative Self-Distillation (NSD) is a new framework for improving large language models by encouraging them to diverge from their own flawed reasoning rather than imitate privileged solutions. Unlike On-Policy Self-Distillation, which can suppress uncertainty and exploratory behavior, NSD generates a question‑specific negative condition (e.g., a careless reasoner) and uses a dynamic gating mechanism to target only reasoning‑critical tokens for penalization. This approach preserves foundational language capabilities while consistently outperforming OPSD and other label‑free self‑bootstrapping reinforcement learning baselines.

By Rongcan Pei, Zhepei Wei, Shuyao Xu, Xinyu Zhu, Wei-Lin Chen, Yu Meng
arXiv Machine Learning
Sep 11

Optimizing Three Critical Factors for Practical and Effective OOD Detection Fine-Tuning

The paper introduces a framework for out-of-distribution (OOD) detection that addresses the trade‑off between detection performance and classification accuracy caused by fine‑tuning with auxiliary outlier data. It optimizes three factors—model reminder, data sampling, and representation learning—by proposing Self‑Knowledge Distillation to preserve accuracy, Semi‑hard Outlier Sampling to enhance detection with minimal data, and Outlier‑aware Supervised Contrastive Learning to improve ID‑OOD separability. The combined approach yields cumulative gains, outperforming existing methods on diverse benchmarks, especially in long‑tailed scenarios, and offers a robust baseline for real‑world OOD detection.

By Hyunjun Choi, JaeHo Chung, Hawook Jeong
arXiv Machine Learning
Sep 11

HERALD: High-Fidelity Exemplar Retrieval with Adaptive Landmark Distillation for Heterophily-Aware Graph Condensation

HERALD is a new gradient‑free graph condensation framework that adapts node scoring and feature selection to a graph’s heterophily level. It selects features using a joint Fisher‑discriminability and activation‑density criterion, and scores nodes with a weighted combination of prototype representativeness, decision‑boundary proximity, and Local Intrinsic Dimensionality, where the weights depend on the heterophily ratio. The selected nodes are assembled into a condensed subgraph via score‑ordered BFS expansion, Personalized PageRank pruning, and class rebalancing, achieving comparable storage to BONSAI and outperforming state‑of‑the‑art condensers on heterophilic graphs while remaining competitive on homophilic ones across multiple GNN architectures.

By Sujan Chakraborty, Priyanka Saha, Saptarshi Bej
arXiv Machine Learning
Sep 11

A Two-Mirror Faceted Projection System for EUV Lithography

The paper proposes an all‑reflective two‑mirror projection system for EUV lithography that achieves a 4× demagnification at a numerical aperture close to unity (NA≈0.993). Unlike conventional EUV objectives that use 6–10 aspheric mirrors and have <15 % throughput, the design uses a fixed two‑reflection path for each accepted diffraction order, retaining 50–60 % of the power and eliminating order‑dependent phase shifts. The authors optimize 30‑bilayer Bragg coatings for each mirror facet, formulate a 3‑D vector model for a two‑dimensionally periodic mask, and use inverse lithography with a differentiable modal solver to demonstrate simulated sub‑10‑nm aerial images with resolved peaks up to 5 nm defocus.

By Vasiliy A. Es'kin, Egor V. Ivanov, Olga V. Martynova
arXiv Machine Learning
Sep 11

FluxMoE: Decoupling Expert Residency for High-Performance MoE Serving

FluxMoE introduces an expert paging system that decouples Mixture-of-Experts (MoE) model experts from permanent GPU residency, allowing dynamic adaptation to available memory. By combining PagedTensor, a bandwidth‑balanced memory hierarchy, and a budget‑aware residency planner, FluxMoE streams expert weights on demand while keeping computations on GPUs. Experiments on GLM‑4.5 and Mixtral‑8×7B‑Instruct show significant throughput gains and reduced time‑per‑output‑token compared to existing inference engines, without compromising model quality.

By Qingxiu Liu, Yongchao He, Runhan Jiang, Zion Wang, Bohan Zhao, Mi Zhang, Patrick P. C. Lee
arXiv Computer Vision
Sep 11

Gaussian Belief Propagation Network for Depth Completion

The paper introduces the Gaussian Belief Propagation Network (GBPN) for depth completion, a hybrid framework that combines deep learning with probabilistic graphical models. GBPN constructs a scene‑specific Markov Random Field via a Graphical Model Construction Network, then infers dense depth distributions using Gaussian Belief Propagation with a serial & parallel message passing scheme. Experiments show GBPN achieves state‑of‑the‑art performance on NYUv2 and KITTI, demonstrating robustness and generalizability across different sparsity levels and patterns.

By Jie Tang, Pingping Xie, Jian Li, Ping Tan
arXiv Machine Learning
Sep 11

PitchFlower: A flow-based neural audio codec with pitch controllability

PitchFlower is a flow‑based neural audio codec that offers explicit pitch controllability by flattening and randomly shifting F0 contours during training while conditioning on the true F0 to reconstruct the original audio. A vector‑quantization bottleneck blocks pitch recovery, and a flow‑based decoder produces high‑quality audio. Experiments demonstrate that PitchFlower matches DSP baselines in pitch accuracy, surpasses state‑of‑the‑art neural codecs in audio quality, and remains robust even when trained on WORLD‑transformed audio, effectively removing vocoder artifacts.

By Diego Torres, Axel Roebel, Nicolas Obin
arXiv Machine Learning
Sep 11

Phase-Decoupled, Model-Calibrated Power Control for Disaggregated LLM Serving

The paper introduces a phase‑decoupled, model‑calibrated power controller for disaggregated large‑language‑model (LLM) serving, addressing the mismatch between GPU power settings and the distinct hardware regimes of prefill and decode stages. By calibrating separate power caps for each lane based on measured throughput‑latency cliffs, the authors achieve a 20.4% increase in tokens per joule with only a 3.5% rise in mean end‑to‑end latency on an 8‑node B200 cluster, outperforming NVIDIA’s Max‑Q profile. The approach also demonstrates consistent meeting of ITL‑p99 service‑level objectives across multiple MoE models and yields a 32.3% electricity savings over a sustained three‑day run.

By Jae Gon Kim, Donghoon Yoo, Hanyul Ryu, Sungho Ha, Juyeon Lee, Soojung Ryu
arXiv Computer Vision
Sep 11

BridgeMatch: Conditional Transport Bridges in Matching Matrix Space for 3D Deformable Registration

BridgeMatch is a two‑stage generative solver that preserves the full soft matching matrix for 3D deformable registration. Stage I uses denoising diffusion to estimate a global matching matrix at a coarse resolution, then lifts it to high resolution while maintaining hierarchy and rank constraints. Stage II refines this lifted matrix via a conditional transport bridge, implemented with either a deterministic Flow Matching ODE or a stochastic Brownian‑bridge SDE, and demonstrates improved correspondence accuracy and registration performance on 4DMatch, 4DLoMatch, CAPE, and DeepDeform datasets, especially in low‑overlap scenarios.

By Qianliang Wu, Haobo Jiang, Guangwei Gao, Shuo Chen, Jin Xie, Jian Yang, Yaqing Ding
arXiv Machine Learning
Sep 11

Longitudinal Risk Prediction in Mammography with Privileged History Distillation

The paper introduces SEM‑HD, a framework that leverages longitudinal mammography history as privileged information during training to improve risk prediction while requiring only a single current exam at inference. By having a student model predict latent representations of past visits and using teacher supervision from actual longitudinal data, SEM‑HD preserves temporal modeling benefits without needing prior exams at deployment. Experiments on three cohorts and two backbone architectures show consistent gains in long‑horizon AUC and pAUC, especially in low false‑positive‑rate regions, and recover much of the performance gap to full‑history models.

By Banafsheh Karimian, Soufiane Belharbi, Alexis Guichemerre, Luke McCaffrey, Mohammadhadi Shateri, Eric Granger
arXiv Machine Learning
Sep 11

A Dataset and Model for Imputing Water Surface Elevation on a Large and Extremely Sparse Spatiotemporal Graph

The paper introduces AmazonSWE, a dataset covering over 19,000 river sections in the Amazon basin for 10 years (2016‑2026) that integrates satellite altimetry, including SWOT, to enable large‑scale spatiotemporal graph imputation. The dataset is extremely sparse—fewer than 1% of sections are observed daily—and features a directed acyclic river topology that is larger and structurally distinct from existing benchmarks. The authors demonstrate that conventional imputation methods struggle with this topology, scale, and sparsity, and propose a bidirectional selective state‑space model that outperforms prior approaches, reducing RMSE against in‑situ gauges by 18‑39% and providing predictions for every river section. whyItMatters":"AmazonSWE offers a novel, real‑world use case that could improve flood forecasting and water resource management by enabling more accurate and comprehensive water surface elevation estimates across a vast, sparsely monitored river network."

By Ruben Cartuyvels, Karim Douch, Gabriele Bertoli, Mounia El Baz, Artemis Vrettou, S\'ebastien Lef\`evre, Diego Fernandez Prieto
arXiv Computer Vision
Sep 11

IMLE-VLA: Fast Single-Step Action Generation for Vision-Language-Action Policies

IMLE‑VLA replaces the iterative action head in vision‑language‑action policies with a single‑step conditional generator trained via conditional Implicit Maximum Likelihood Estimation (cIMLE). This eliminates multi‑step sampling, boosting inference frequency by 3.67× (55 Hz vs. 15 Hz) and achieving the highest average success rate (98.0 %) on the 40‑task LIBERO benchmark while maintaining robustness under perturbations. Real‑world tests on a Franka Emika Panda show smoother, faster motions and a 3.9×–6.6× reduction in inference time per episode.

By Kian Hosseinkhani (Simon Fraser University), Qinhe Peng (University of Pennsylvania), George Shramko (Simon Fraser University), Mehran Aghabozorgi (Simon Fraser University), Jianing Qian (University of Pennsylvania), Tristan Engst (Simon Fraser University), Alireza Moazeni (Simon Fraser University), Dinesh Jayaraman (University of Pennsylvania), Ke Li (Simon Fraser University, Canada CIFAR AI Chair)
arXiv Machine Learning
Sep 11

Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data

The paper investigates how data repetition affects Mixture-of-Experts (MoE) language models compared to dense Transformers. Across models from 80 M to 1 B active parameters, MoEs degrade more quickly as data is repeated, with performance dropping significantly beyond 4× repetition and overtaking dense models only when strong regularization is applied. The study also identifies routing stabilization and expert specialization as key factors in MoE overfitting, and explores regularization techniques that can partially mitigate this issue.

By Atindra Jha, Margaret Li, Jure Leskovec, Percy Liang, Luke Zettlemoyer
arXiv Computation and Language
Sep 11

FlexComp: One Model for Every Ratio in Context Compression

FlexComp is a framework that allows a single model to perform context compression at any desired ratio, unlike existing methods that require separate models for each fixed ratio. It achieves this by sampling a memory budget during training and selecting the appropriate budget at inference time using either confidence-based cascade routing or a lightweight learned predictor. Experiments on ICAE, 500xCompressor, and SAC show that FlexComp matches the performance of specialized fixed-ratio models while enabling high compression rates and improving decoding throughput.

By Kaiyan Zhao, Zhongtao Miao, Akiko Aizawa, Yoshimasa Tsuruoka
arXiv Machine Learning
Sep 11

Sparsity Regularized and Robust Mean Variance Portfolio Selection Under Ellipsoidal Uncertainty

The paper studies mean‑variance portfolio selection with an β0 penalty to encourage sparse asset allocations. It incorporates uncertainty in the mean return vector via an ellipsoidal uncertainty set, leading to a robust sparse optimization framework. The authors analyze the structure of local and global minimizers, develop a branch‑and‑bound algorithm with a novel pruning rule, and show through computational experiments that their method is effective and competitive with existing solvers.

By Deniz Akkaya, Emre Can Yayla, Buse \c{S}en, Mustafa \c{C}. P{\i}nar
arXiv Computation and Language
Sep 11

A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and Distillation

The paper proposes a new method for training large language models to handle long-context reasoning by combining Group Relative Policy Optimization (GRPO) with on‑policy distillation (OPD). It introduces a synthetic multilingual dataset called LongBlocks that tests multi‑hop reasoning, contextual grounding, and long‑form generation. Experiments show that the combined approach outperforms either GRPO or OPD alone while maintaining short‑context performance.

By Miguel Moura Ramos, Duarte M. Alves, Andr\'e F. T. Martins
arXiv Machine Learning
Sep 11

Phases in a class of associative memories via hidden neurons

The paper investigates associative memory in a bipartite Hopfield–Krotov architecture, termed class H, where hidden neurons serve as the retrieval order parameter. Using the replica method, it derives replica‑symmetric phase diagrams and closed‑form capacities for polynomial load, showing that crosstalk statistics are similar for Ising and spherical visible neurons. With a softmax hidden layer, the load becomes exponential, mapping the thermodynamics onto a random‑energy‑model that exhibits paramagnetic, condensed, and frozen phases, and revealing that heating destabilizes retrieval through quantized attention reassignments while Gaussian patterns remain metastable at all loads.

By Toshihiro Ota, Masato Taki
arXiv Computation and Language
Sep 11

Beyond Prompting: Efficient and Robust Contextual Biasing for Speech LLMs via Logit-Space Integration (LOGIC)

The paper introduces LOGIC (Logit‑Space Integration for Contextual Biasing), a new framework that injects contextual entity information directly into the decoding layer of Speech Large Language Models, bypassing the limitations of prompt‑based methods. LOGIC operates with constant‑time complexity regardless of the size of the entity list, and experiments with the Phi‑4‑MM model across 11 multilingual locales show an average 9% relative reduction in Entity WER while adding only a 0.30% increase in False Alarm Rate.

By Peidong Wang, Jian Xue, Jinyu Li