Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,914 stories · RSS feed

arXiv AI
Sep 17

Visual Perception Engine: Fast and Flexible Multi-Head Inference for Robotic Vision Tasks

Visual Perception Engine (VPEngine) is a modular framework that enables efficient GPU usage for robotic vision tasks by sharing a foundation model backbone across multiple specialized task heads. It eliminates redundant feature extraction, supports dynamic task prioritization, and achieves up to 3× speedup over sequential execution. The open‑source Python implementation, with ROS2 C++ bindings, delivers real‑time performance (≥50 Hz) on NVIDIA Jetson Orin AGX using TensorRT‑optimized models.

By Jakub {\L}ucki, Jonathan Becktor, Georgios Georgakis, Rob Royce, Shehryar Khattak
arXiv AI
Sep 17

Ultralytics YOLO Evolution: An Overview of YOLO27, YOLO26, YOLO11, YOLOv8, and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition

This paper provides a detailed overview of the Ultralytics YOLO family from YOLOv5 to YOLO27, highlighting key architectural changes, benchmarking results, and deployment considerations. It discusses the evolution of each version—YOLO27’s scale‑adaptive dual architecture, YOLO26’s loss and optimization refinements, YOLO11’s efficiency focus, YOLOv8’s anchor‑free detection, and YOLOv5’s modular ecosystem—alongside performance metrics on COCO and latency on TensorRT. The review also surveys applications in robotics, agriculture, surveillance, and manufacturing, and outlines future challenges such as dense scene handling, CNN‑Transformer integration, and hardware‑aware optimization.

By Ranjan Sapkota, Manoj Karkee
arXiv AI
Sep 17

VISTA: Validation-Informed Trajectory Adaptation via Self-Distillation

VISTA is an online self‑distillation framework that enforces consistency along a deep learning model’s optimization trajectory. It uses a validation‑informed Marginal Coverage score to identify earlier model states—called expert anchors—that retain specialized competence over distinct data regions. By integrating a coverage‑weighted ensemble of these anchors during training, VISTA regularizes the loss landscape, preserves learned knowledge, and improves robustness and generalization while cutting storage overhead by 90%.

By Eli Corn, Daphna Weinshall
arXiv Computation and Language
Sep 17

Size Matters: Foundation Model for Czech HTML documents

The paper introduces HTML‑LM, a 154‑million‑parameter foundation model designed for Czech HTML documents. It leverages HTML‑aware training and a ModernBERT architecture, trained on 100 million web pages with objectives such as masked language modeling, bag‑of‑words prediction, and contrastive distillation from larger language models. HTML‑LM achieves state‑of‑the‑art performance on classification and regression tasks in the Czech Internet domain, outperforms larger encoders and small LLMs, and is deployed in production to process thousands of web documents per second.

By Martin Dvo\v{r}\'ak, V\'it Tlusto\v{s}, Artyom Voronin, Martin Habrovec, Kate\v{r}ina Podlesn\'a, Barbora Ri\v{s}ov\'a, Josef Von\'a\v{s}ek
arXiv Machine Learning
Sep 17

Robust and Efficient AI Frameworks for Scalable Material Design and Property Prediction

The thesis presents AI frameworks that accelerate crystalline materials discovery by tackling both crystal property prediction and crystal structure generation. It introduces CrysXPP, CrysGNN, and CrysMMNet for efficient, data‑sparse property prediction using graph autoencoding, self‑supervised pretraining, and multimodal learning. For generation, TGDMat is a text‑guided diffusion model that jointly learns lattice parameters, atomic types, and coordinates, enabling valid, stable, and conditionally generated periodic materials.

By Kishalay Das
arXiv Machine Learning
Sep 17

Gradient Descent with Stochastic Subspaces via Persistence of Memory

The paper introduces a novel technique called "persistence of memory" to enhance stochastic subspace methods for large‑scale optimisation. By using a weakly correlated guidance vector that is refreshed only at wide intervals, the method provides a structured direction for random subspace descent. The authors demonstrate that this guidance can be efficiently computed in sparse or minibatch settings and present the first theoretical analysis of classical SSD methods for sparse functions, showing alignment with low‑lying Hessian eigenvectors near the optimum.

By Subhroshekhar Ghosh, Clement Z. Q. Ng, Pierre-Louis Poirion, Akiko Takeda
arXiv Machine Learning
Sep 17

Reinforcement Learning for Real-Time Vision-Language-Action Policies

The paper presents Real‑Time EXPO‑FT, a reinforcement learning framework that fine‑tunes large Vision‑Language‑Action models for real‑time robotic control. It separates slow, expressive action generation from fast, reactive edits, allowing a lightweight policy to adjust actions based on the latest observation. Experiments on the Kinetix benchmark and four dynamic real‑world tasks show that Real‑Time EXPO‑FT achieves superior performance, improving policy success rates from 42% to 97% with only ten minutes of online data and no human intervention.

By Perry Dong, Kuo-Han Hung, Dorsa Sadigh, Chelsea Finn
arXiv AI
Sep 17

Performance and Complexity Trade-off Optimization of Speech Models During Training

The paper introduces a reparameterization technique that injects feature noise to jointly optimize speech model performance and computational complexity during training. Unlike traditional pruning, this method dynamically adjusts model size for a desired performance‑complexity trade‑off without heuristic weight removal. The authors validate their approach with a synthetic example and two real‑world applications—voice activity detection and audio anti‑spoofing—providing publicly available code for further research.

By Esteban G\'omez, Tom B\"ackstr\"om
arXiv Machine Learning
Sep 17

Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate

The paper introduces DASH-Q, a post‑training quantization method that uses a diagonal Hessian approximation and iterative weighted least squares to reduce noise in curvature estimates. By discarding noisy cross‑channel dependencies, DASH‑Q preserves salient feature power and achieves superior performance in ultra low‑bit quantization. Across five large language models, it improves zero‑shot accuracy by an average of 7.01% and up to 14.01% over the strongest baselines, even with very small calibration datasets.

By Jaemin Kim, Sungkyun Kim, Junyeol Lee, Jiwon Seo
arXiv Computer Vision
Sep 17

PhysVGGT: Feed-Forward Dense Physical Property Estimation from A Single Image

PhysVGGT is a feed‑forward model that predicts dense maps of friction coefficient, Shore hardness, Young's modulus, and density, along with object‑level mass, from a single RGB image in one forward pass. It treats physical property estimation as a dense per‑pixel prediction problem, using a visual geometry transformer to extract geometry‑aware tokens and separate dense and global prediction branches. A scalable pseudo‑label generation pipeline enables large‑scale weakly supervised training, and the model achieves state‑of‑the‑art performance on the ABO‑500 dataset while running 27× faster than previous methods.

By Sneha Paul, Guile Wu, Bingbing Liu, Dongfeng Bai