Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,731 stories · RSS feed

arXiv Machine Learning
Sep 24

Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery

The paper introduces a method to prune six layers from the encoder of OpenAI’s Whisper ASR model, reducing the encoder stack by 18.5% without requiring custom inference code. Layers are selected based on their minimal impact on Word Error Rate when removed. After pruning, the model’s WER rises from 18.2% to 21.9%, but distillation with unlabeled monolingual speech data lowers it to 20.1%. "whyItMatters":"The approach offers a straightforward way to accelerate Whisper inference by simplifying the encoder while maintaining acceptable accuracy, and the released code and model enable immediate adoption by the community."

By Rasmus Aagaard, Nicki Skafte Detlefsen
arXiv Machine Learning
Sep 24

RAMP: Robust Adaptive Mixed-Precision Quantization for Edge CPU Vision Models

The paper introduces RAMP, a method for robust adaptive mixed‑precision quantization of vision models on edge CPUs. It evaluates 13 sensitivity metrics across four neural networks, finding that Jensen‑Shannon Divergence consistently identifies layers that can be safely quantized. Using K‑Means clustering on these metrics, RAMP achieves near‑lossless accuracy with an average 1.81× speed‑up, while cautioning against excluding low‑speed‑up layers that can fragment the computational graph.

By David Poblaci\'on-Criado, Dario Garcia-Gasulla, Eduardo Quinones
arXiv Machine Learning
Sep 24

Predicting Quantization Price for Selecting PTQ Configurations Before Deployment

The paper proposes a method for selecting post‑training quantization (PTQ) configurations before deployment by treating each admissible layer configuration as an error generator with an associated deployment cost. It introduces a priced layer‑output error framework that uses the covariance of layer outputs and the full‑precision model’s curvature to compute a price for each configuration. This approach replaces traditional reconstruction or Hessian‑based scores with a unified, cost‑aware selector that can calibrate and budget PTQ settings efficiently.

By Junbin Qiu, Jian Mu, Weitong Zhang, Yao Shu
arXiv Machine Learning
Sep 24

Integrated Multivariate Segmentation Tree for Heterogeneous Credit Data Analysis in Small- and Medium-Sized Enterprises

The paper introduces the Integrated Multivariate Segmentation Tree (IMST), a new framework that combines financial data and textual information for credit evaluation of small- and medium-sized enterprises. IMST transforms text into numerical matrices via matrix factorization, selects key financial features with Lasso regression, and builds a multivariate segmentation tree using Gini or entropy with weakest-link pruning. Experiments on 1,428 Chinese SMEs show an 88.9% accuracy, outperforming baseline decision trees, SVMs, and neural networks while offering better interpretability and computational efficiency.

By Lu Han, Xiuying Wang
arXiv Machine Learning
Sep 24

Statistical Properties of Deep Neural Networks with Dependent Data

The paper develops theory for deep neural network (DNN) estimators under dependent data. It establishes nonasymptotic probability bounds on the theoretical and empirical ∼2-errors of nonparametric sieve estimators for a general class of estimation problems with possibly nonstationary β-mixing data in unbounded sets. The theory is then applied to fully connected and convolutional DNN estimators without weight bounds or sparsity restrictions, deriving results for H"older smooth functions under nonstationary, subgaussian, β-mixing data with exponential or polynomial decay, and achieving the nonparametric minimax rate up to logarithmic factors in several regression settings.

By Chad Brown
arXiv Computer Vision
Sep 24

SatUnreal: A High-Precision Synthetic Dataset for Satellite Stereo Matching via Unreal Engine

SatUnreal is a synthetic dataset created with Unreal Engine that offers 10,000 high‑resolution (0.3 m GSD) satellite stereo pairs. It addresses key limitations of existing benchmarks by ensuring physical geometry simulation, spatio‑temporal consistency, topographic diversity, and mathematically precise occlusion masks via a two‑step line‑trace algorithm. Models trained solely on SatUnreal outperform those trained on real datasets when transferred to real‑world benchmarks such as US3D and WHU‑Stereo.

By Han-Gyeol Kim, JaeWan Park, Junmin Park, Darongsae Kwon
arXiv Computer Vision
Sep 24

LiAM-SAM: Lifecycle-Aware Memory for Robust SAM2-Based MOT

LiAM‑SAM is a lifecycle‑aware memory framework designed to improve segmentation‑based multi‑object tracking (MOT) with the SAM2 foundation video model. It addresses three common failure modes—faulty track initiation, memory drift during close interactions, and unreliable re‑identification after occlusion—by introducing contrastive track initiation, motion‑ and geometry‑grounded memory correction, and adaptive context memory. The system achieves state‑of‑the‑art HOTA and IDF1 scores, with ablations showing significant gains in association metrics and a 96% reduction in identity switches.

By Gr\'egoire Francisco, Alessandro D'Amico, Samuele Costantini, Gianpiero Francesca, Lorenzo Garattoni
arXiv Computer Vision
Sep 24

OD3: Optimization-free Dataset Distillation for Object Detection

OD3 introduces an optimization‑free dataset distillation framework tailored for object detection. The method first iteratively places object instances in synthesized images, then screens candidates with a pre‑trained observer model to discard low‑confidence objects. Applied to MS COCO and PASCAL VOC, OD3 achieves compression ratios from 0.25% to 5% and surpasses previous detection‑focused distillation methods by over 14% on COCO mAP50 at a 1.0% compression ratio.

By Salwa K. Al Khatib, Ahmed ElHagry, Shitong Shao, Zhiqiang Shen
arXiv Machine Learning
Sep 24

CRISP: Scalable Importance-Stratified Coresets for Imbalanced Tabular Learning

CRISP (Coreset Reduction via Importance-Stratified Pruning) is a linear-time method that reduces negative-class examples in highly imbalanced tabular datasets by allocating a budget across quantile strata of a proxy-model score and using sample weights to correct for unequal inclusion probabilities. On a production fraud dataset, CRISP cuts the training set from 25 M to about 1.70 M rows (a 93.2% reduction) while preserving 99.7% of the full-data Average Precision. In public benchmarks such as CriteoPrivateAds, CRISP consistently achieves the highest mean Average Precision across a range of majority reductions, with ablation studies highlighting budget allocation and inverse-propensity weighting as key contributors to its performance.

By Hardhik Mohanty, Indrayana Rustandi, Mohamadreza Sheibani
arXiv Machine Learning
Sep 24

GeoRVQ: Decoder-aware geometry for residual-token prediction in physiological signals

GeoRVQ introduces a coarse‑to‑fine masked token model that incorporates decoder‑induced geometry into residual‑vector‑quantized physiological waveforms. By using geometry‑aware soft targets and expected distortion, the model improves exact token accuracy from 0.133 to 0.143, reduces decoded distance from 0.606 to 0.393, and raises R‑peak F1 from 0.784 to 0.837 on datasets such as MIMIC‑IV Waveform, VitalDB, and CODE‑15%. The approach demonstrates that decoder‑aware objectives can enhance waveform and event preservation without a large increase in token accuracy.

By Bo Cui, Yaowen Zhang
arXiv Machine Learning
Sep 24

WTF?! Simulation-Free Reinforcement Learning with Wasserstein-Tilted Flow Maps

The paper introduces Wasserstein‑Tilted Flow Maps (WTF), a simulation‑free reinforcement learning method that fine‑tunes pre‑trained flow‑based generative models by adding an optimal transport regularizer derived from the model’s drift. Unlike traditional KL‑reward tilting, WTF transports individual samples toward higher reward, framing the problem as a deterministic optimal control task on the flow map. Experiments on ImageNet‑256 and text‑to‑image demonstrate that WTF achieves higher reward and comparable or better diversity while reducing training compute by up to 280×.

By Abbas Mammadov, Jerry Y. Huang, Justin Lin, Partha Kaushik, Sheel Shah, Kartik Nair, Yee Whye Teh, Nicholas M. Boffi
arXiv Machine Learning
Sep 24

Graph Learning with Spectral Connectivity Priors for Scarce Data

The paper introduces Spectral Connectivity-Regularized Graph Learning (SCoGL), a method for learning sparse graphs from limited data by incorporating Laplacian spectral priors that promote global connectivity. SCoGL extends the graphical lasso objective with a connectivity prior derived from Laplacian eigenvalues and uses projected gradient descent with Armijo backtracking for optimization. Experiments demonstrate that SCoGL improves graph recovery and enhances downstream tasks such as graph signal denoising when observations are scarce.

By Mingxiao Liu (Tsinghua University, China), Bahar Oveisgharan (York University, Canada), Bingyan Zou (Tsinghua University, China), Gene Cheung (York University, Canada), H. Vicky Zhao (Tsinghua University, China), Feifei Gao (Tsinghua University, China)
arXiv Machine Learning
Sep 24

MENO: Memory-Efficient Neural Operator

The paper introduces MENO, a Memory‑Efficient Neural Operator designed for solving partial differential equations (PDEs). MENO leverages a Manifold Function Encoder to achieve a small memory footprint that does not depend on data resolution, enabling faster training and potential scalability to large models. It accepts PDE inputs of arbitrary form—including different geometric domains and discretizations—allowing cross‑geometry scenarios, and demonstrates strong generalization with superior accuracy on most tested benchmarks.

By Shengyang Xu, Weijun Zhang, Jun Hu, Pengzhan Jin
arXiv Machine Learning
Sep 24

Pheno-GS: Phenoscape-scale Geodesic Sinkhorn

Pheno-GS is a new method for computing scalable, geometry-aware optimal transport distances between large patient cohorts of single‑cell data. It uses graph connectivity regularization, an unbalanced OT formulation with KL marginal penalties, and a batched matrix algorithm that dramatically speeds up pairwise distance calculations. The authors validate the approach on synthetic benchmarks and a CyTOF perturbation dataset.

By Alistair Wilkinson, Christopher J. Tape, Smita Krishnaswamy
arXiv Machine Learning
Sep 24

hyperbolix: Hyperbolic Deep Learning in JAX

hyperbolix is an open‑source library for hyperbolic deep learning in JAX, built on Flax NNX. It provides six manifolds—including Euclidean, Poincaré ball, hyperboloid, κ‑stereographic, mixed‑curvature product, and proper velocity space—through a common interface, and implements a wide range of layer families (linear, convolution, attention, normalization, positional encoding, regression, vector quantization). The library also supplies Riemannian optimizers, wrapped distributions, dimensionality‑reduction techniques, and precision‑tested operations that replace numerically unstable formulas on the hyperboloid, ensuring accurate float32 computations at large distances.

By Timo Klein, Thomas Lang, Yllka Velaj, Sebastian Tschiatschek
arXiv Machine Learning
Sep 24

Minimal-Norm Univariate Two-Layer ReLU Classification: Exact Solutions and Global Optimality with Skip Connections

The paper investigates minimal‑norm interpolation and λ2‑regularized logistic‑loss minimization for binary classification using univariate two‑layer ReLU networks. It provides exact geometric characterizations of optimal classifiers, showing that unpenalized hidden‑layer biases yield continuous piecewise‑affine functions that tightly follow label switches, while penalized biases produce a unique, sparsest classifier with a single kink per same‑label segment. Adding a free affine skip connection does not change these function‑space solutions but guarantees that every KKT point becomes globally optimal, eliminating suboptimal KKT points that can arise without the skip connection.

By Karolina Drabik, Ben Lewis, Antoni Puch, Etienne Boursier, Piotr Hofman, Matthias Englert, Ranko Lazi\'c
arXiv Machine Learning
Sep 24

The Computational Value of Sensory-Aligned Receptive Fields Depends on Neuronal Expressivity

The study investigates whether sensory-aligned receptive fields provide computational benefits beyond mere resource efficiency in recurrent networks of Expressive Leaky Memory neurons. Across auditory and event-based visual classification tasks, receptive fields aligned with task-relevant sensory coordinates improve test accuracy compared to budget-matched random fields, but this advantage disappears when coordinates are scrambled or irrelevant. The benefit diminishes as neuronal expressivity increases, and generic synaptic sparsity regularization only partially recovers performance, indicating that structured receptive fields act as a computational prior beyond sparsity alone.

By Agnese Adorante, Aaron Spieler, Anna Levina