Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

6,032 stories · RSS feed

arXiv Computer Vision
Sep 11

Uncertainty DMD: Restoring Diversity in Few-Step Autoregressive Video Distillation

The paper introduces Uncertainty DMD, a lightweight framework that injects uncertainty into few-step autoregressive video distillation to counteract diversity collapse. By perturbing the first chunk’s timestep and employing a stochastic cache-writing mechanism for subsequent chunks, the method restores stochasticity without altering the model architecture. Experiments demonstrate consistent improvements in video diversity and motion dynamics while preserving visual quality.

By Zixuan Duan, Xunzhi Xiang, Yabo Chen, Xin Zhang, Changhan Liu, Haibin Huang, Chi Zhang, Qi Fan, Xuelong Li
arXiv Machine Learning
Sep 11

EMMI: Edge Multi-Modal Intelligence for Communication-Efficient MLLM Inference via Fused Representation Compression

The paper introduces EMMI, a framework that enables communication‑efficient inference of multimodal large language models (MLLMs) on edge devices. EMMI encodes each sensor modality separately, fuses the representations, and compresses them into a compact latent vector that is transmitted to a server for high‑capacity reasoning. Experiments on a multimodal benchmark show that EMMI can cut the communication payload by 32× while keeping accuracy comparable, achieving up to a 3.4× reduction in end‑to‑end inference latency under bandwidth‑constrained conditions.

By Motahare Mounesan, Irfan Khan
arXiv Computation and Language
Sep 11

KuaiRP Series Role-playing Models Technical Report

The paper presents the KuaiRP series of role‑playing models, detailing a multi‑stage training pipeline that balances deep domain knowledge injection with the preservation of general agent capabilities. The approach includes a standardized character template, a supervised fine‑tuning (SFT) data pipeline, a rule‑based reward function for reinforcement learning, and a novel two‑stage on‑policy distillation (OPD) with cumulative‑divergence decay (CDD) to recover general skills. Experimental results show that the models achieve state‑of‑the‑art role‑playing fidelity in target domains while maintaining low deployment costs.

By Yipeng Wang, Ziwei Zhang, Jiahui Zhang, Qi Gan, Kai Sheng
arXiv Machine Learning
Sep 11

BiHDTrans: binary hyperdimensional transformer for efficient multivariate time series classification

BiHDTrans is a neurosymbolic binary hyperdimensional transformer that merges self‑attention with hyperdimensional computing to classify multivariate time series efficiently. It surpasses existing HD models by at least 14.47% and binary transformers by 6.67% on average, while an FPGA‑accelerated implementation reduces inference latency 39.4× compared to state‑of‑the‑art binary transformers. Even with a 64% reduction in hyperspace dimensionality, BiHDTrans remains competitive, achieving 1–2% higher accuracy with 4.4× smaller model size and nearly 50% lower latency than the full‑dimensional baseline.

By Jingtao Zhang, Yi Liu, Qi Shen, Changhong Wang
arXiv Machine Learning
Sep 11

Longitudinal Risk Prediction in Mammography with Privileged History Distillation

The paper introduces SEM‑HD, a framework that leverages longitudinal mammography history as privileged information during training to improve risk prediction while requiring only a single current exam at inference. By having a student model predict latent representations of past visits and using teacher supervision from actual longitudinal data, SEM‑HD preserves temporal modeling benefits without needing prior exams at deployment. Experiments on three cohorts and two backbone architectures show consistent gains in long‑horizon AUC and pAUC, especially in low false‑positive‑rate regions, and recover much of the performance gap to full‑history models.

By Banafsheh Karimian, Soufiane Belharbi, Alexis Guichemerre, Luke McCaffrey, Mohammadhadi Shateri, Eric Granger
arXiv Machine Learning
Sep 11

PitchFlower: A flow-based neural audio codec with pitch controllability

PitchFlower is a flow‑based neural audio codec that offers explicit pitch controllability by flattening and randomly shifting F0 contours during training while conditioning on the true F0 to reconstruct the original audio. A vector‑quantization bottleneck blocks pitch recovery, and a flow‑based decoder produces high‑quality audio. Experiments demonstrate that PitchFlower matches DSP baselines in pitch accuracy, surpasses state‑of‑the‑art neural codecs in audio quality, and remains robust even when trained on WORLD‑transformed audio, effectively removing vocoder artifacts.

By Diego Torres, Axel Roebel, Nicolas Obin
arXiv Machine Learning
Sep 11

Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data

The paper investigates how data repetition affects Mixture-of-Experts (MoE) language models compared to dense Transformers. Across models from 80 M to 1 B active parameters, MoEs degrade more quickly as data is repeated, with performance dropping significantly beyond 4× repetition and overtaking dense models only when strong regularization is applied. The study also identifies routing stabilization and expert specialization as key factors in MoE overfitting, and explores regularization techniques that can partially mitigate this issue.

By Atindra Jha, Margaret Li, Jure Leskovec, Percy Liang, Luke Zettlemoyer
arXiv Computation and Language
Sep 11

A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and Distillation

The paper proposes a new method for training large language models to handle long-context reasoning by combining Group Relative Policy Optimization (GRPO) with on‑policy distillation (OPD). It introduces a synthetic multilingual dataset called LongBlocks that tests multi‑hop reasoning, contextual grounding, and long‑form generation. Experiments show that the combined approach outperforms either GRPO or OPD alone while maintaining short‑context performance.

By Miguel Moura Ramos, Duarte M. Alves, Andr\'e F. T. Martins
arXiv Computer Vision
Sep 11

Does YOLO26 Truly Offer Advantages Over Its Predecessors for Edge Deployment? A Benchmark Study in Aquaculture

The paper evaluates the new YOLO26 architecture, which offers NMS-free end-to-end inference and is tailored for CPU-based edge devices, against three earlier Ultralytics models (YOLOv5u, YOLOv8, and YOLO11) in aquaculture fish mortality detection. Across nano, small, and medium scales, all models achieved similar detection accuracy on a full dataset, but differences emerged in data efficiency and deployment performance: YOLOv8 reached 90% mAP50 with only 400 images, while YOLO26 variants needed 1,000 images; YOLO26n was fastest on a Raspberry Pi 5 (7.51 FPS), whereas YOLOv5mu led on CPU-based hardware. The study concludes that architectural novelty alone does not dictate suitability for edge AI in aquaculture; training data size, target hardware, and inference needs must be jointly considered.

By Rakesh Ranjan, Gajanan S. Kothawade, Kata Sharrer, Scott Tsukuda, Christopher Good
arXiv Machine Learning
Sep 11

Positional task conditioning for scalable defect detection across product families in large product catalogs

The paper presents a method called Positional Task Conditioning (PTC) to improve defect detection in large product catalogs. By breaking detection into focused sub‑tasks and reinforcing task identity at prompt boundaries, PTC reduces context length and isolates error types, boosting F1 scores from 52% to 87%. The approach outperforms rationale‑based distillation across multiple models, achieving near‑state‑of‑the‑art performance at up to 98% lower cost and is deployed in several countries handling over 10 million product families.

By Soham Satyadharma, Gabriel Roccabruna, Suleiman A. Khan
arXiv Computer Vision
Sep 11

Gaussian Belief Propagation Network for Depth Completion

The paper introduces the Gaussian Belief Propagation Network (GBPN) for depth completion, a hybrid framework that combines deep learning with probabilistic graphical models. GBPN constructs a scene‑specific Markov Random Field via a Graphical Model Construction Network, then infers dense depth distributions using Gaussian Belief Propagation with a serial & parallel message passing scheme. Experiments show GBPN achieves state‑of‑the‑art performance on NYUv2 and KITTI, demonstrating robustness and generalizability across different sparsity levels and patterns.

By Jie Tang, Pingping Xie, Jian Li, Ping Tan
Hugging Face Trending Papers
Sep 10

Sparsity Regularized and Robust Mean Variance Portfolio Selection Under Ellipsoidal Uncertainty

The paper studies mean‑variance portfolio selection using an β0 penalty to encourage sparse asset allocations. It incorporates uncertainty in expected returns via an ellipsoidal set, leading to a robust sparse optimization framework. The authors analyze local and global minimizers, design a branch‑and‑bound algorithm with a novel pruning rule, and show through computational experiments that their method outperforms a mixed‑integer second‑order cone programming solver on real market data.

Hugging Face Trending Papers
Sep 10

Negative Self-Distillation: Learning to Reason by Avoiding Flaws

Negative Self-Distillation (NSD) is a new framework for improving large language models by encouraging them to diverge from their own flawed reasoning rather than imitate privileged solutions. Unlike On-Policy Self-Distillation (OPSD), which can suppress uncertainty and exploratory behavior, NSD generates a question‑specific negative condition (e.g., a careless reasoner) and pushes the student’s distribution away from it. A dynamic gating mechanism isolates reasoning‑critical tokens so that only behavioral flaws are penalized, preserving linguistic capabilities, and empirical results show NSD consistently outperforms OPSD and other label‑free self‑bootstrapping RL baselines.

Hugging Face Trending Papers
Sep 10

ZipCodec: Ultra-Low-Frame-Rate Streaming Speech Coding

ZipCodec is a streaming neural speech codec that operates at an ultra‑low frame rate of 6.25 Hz and a bitrate of 0.80 kbps, achieving a theoretical latency of 160 ms. It leverages large‑scale WavLM distillation, a redesigned transformer architecture, a scalar spherical quantizer, and a latency‑aware streaming decoder. Experiments demonstrate that ZipCodec outperforms existing streaming codecs at comparable bitrates in both reconstruction quality and downstream tasks, while remaining real‑time on a consumer‑grade CPU despite its 842 M parameters.

Hugging Face Trending Papers
Sep 10

A Dataset and Model for Imputing Water Surface Elevation on a Large and Extremely Sparse Spatiotemporal Graph

The paper introduces AmazonSWE, a dataset covering over 19,000 river sections in the Amazon basin for 10 years, combining satellite altimetry and in‑situ gauge data to enable large‑scale spatiotemporal graph imputation. The authors highlight the extreme sparsity of observations—less than 1% of sections per day—and the directed acyclic topology of river networks, which challenge existing imputation methods. They propose a bidirectional selective state‑space model that samples connected subgraphs and uses topology‑aware positional encodings, achieving 18–39% lower RMSE than the current state‑of‑the‑art SWOT‑based approach while providing predictions for all river sections.

Hugging Face Trending Papers
Sep 10

Phase-Decoupled, Model-Calibrated Power Control for Disaggregated LLM Serving

The paper evaluates NVIDIA’s Max‑Q inference profile on a disaggregated B200 GPU system for large language model (LLM) serving, finding modest gains (+8.6% tokens/J) but increased latency (+5.2%). It proposes a phase‑decoupled, model‑calibrated power controller that sets a latency‑guaranteed SM‑clock window for prefill and a calibrated power cap for decode, achieving a 20.4% tokens/J improvement with only a 3.5% latency increase on an 8‑node Qwen3‑Coder‑480B deployment. The approach outperforms vendor profiles on both energy and latency, and demonstrates significant long‑term electricity savings in MoE‑based serving.

arXiv Machine Learning
Sep 10

High-dimensional Linear Bandits with Knapsacks

The paper studies high‑dimensional linear contextual bandits with knapsack constraints (CBwK), aiming to exploit sparsity for tighter regret bounds. It introduces an online hard‑thresholding estimator integrated into a primal‑dual framework, achieving sub‑linear regret that grows only logarithmically with the feature dimension. Under either a diverse‑covariate or margin condition, the regret improves to τ‑dependent rates, and when both hold simultaneously, a dual resolving scheme yields an even tighter bound. The approach also recovers optimal rates for high‑dimensional contextual bandits without knapsacks, and experiments demonstrate its practical effectiveness.

By Wanteng Ma, Dong Xia, Jiashuo Jiang