Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,633 stories · RSS feed

arXiv Machine Learning
3d ago

XS-VLA: Teaching Tiny Vision-Language-Action Models with Spatial Supervision and Demonstration Conditioning

arXiv:2607.04171v4 Announce Type: replace-cross Abstract: How can richer training supervision improve robot control while keeping the deployed policy compact? We present XS-VLA, a staged training fra...

By Iok Tong Lei, Ying Jie Yap, Wei Huang, Qingchen Xie, Qianzhi Li, Yujie Zhang, Xiaolong Liu, Zhidong Deng
arXiv Computer Vision
3d ago

Joint Branch-Space Transform Coding for Diffusion Activation Quantization with Classifier-Free Guidance

arXiv:2610.00930v1 Announce Type: new Abstract: Post-training quantization for diffusion models increasingly exploits timestep, feature, and layer structure. While recent work has begun incorporating...

By Mingrun Jiang, Yuejia Liu, Zishan Shao, Ting Jiang, Qinsi Wang, Hancheng Ye, Yixiao Wang, Rui-Feng Wang, Kangning Cui, Yixuan Chen, Fan Yang, Xiang Cheng, Hai Li, Yiran Chen
arXiv Computer Vision
3d ago

MMVistaReason: Toward Open-Data and Post-Training Recipes for Multimodal Reasoning

arXiv:2610.01352v1 Announce Type: new Abstract: Open multimodal reasoning models have benefited from large-scale reasoning supervision, yet reliable post-training remains challenging due to uneven da...

By Juekai Lin, Honglin Lin, Yuqian Yuan, Xiaolong Wu, Jie Cao, Liang Liang, Yunqi Cao, Yun Zhu, Wenqiao Zhang, Lijun Wu
arXiv Computer Vision
3d ago

DecomVoxel: Harnessing 3D-Native Priors with Guided In-situ Denoising Optimization for Decompositional Scene Reconstruction

DecomVoxel introduces a guided in‑situ denoising optimization that fuses 3D‑native priors with neural scene reconstruction to improve decompositional scene reconstruction. The method employs an epsilon‑based distillation loss for stable latent refinement and adaptive spatial guidance using occupied and vacant anchors with temporal annealing to reduce hallucinations and spatial drift. Experiments on Replica and ScanNet++ demonstrate that DecomVoxel outperforms state‑of‑the‑art approaches while preserving spatial layout, structural fidelity, and style‑consistent texture, yielding high‑quality textured meshes with clean topology.

By Junfeng Ni, Zirui Zhou, Yixin Chen, Yu Liu, Nan Jiang, Zhifei Yang, Song-Chun Zhu, Siyuan Huang
arXiv Computer Vision
3d ago

Dyna-DINO: Efficient ViT Distillation Via Adaptive Representation Anchoring

Dyna‑DINO introduces a curriculum for Vision Transformer (ViT) knowledge distillation that uses the teacher’s intermediate feature maps as progressively harder targets, enabling a student to build foundational representations before tackling higher‑level abstractions. The approach accelerates convergence and improves performance across multiple tasks: on ImageNet‑100 the distilled ViT‑S reaches 90.1% accuracy (+12.24% over baseline), while on ImageNet‑1K it yields +3.9% and +6.09% gains on Oxford and Paris retrieval, +1.93% on semantic segmentation, and notable classification improvements. Additionally, the curriculum reduces training FLOPs by 25.1% and training time by 21% on ImageNet‑100 through early‑stopping of teacher inference.

By Jiaqi Zhang, Ashton Lee, Anthony Wong, John Zou, Sami BuGhanem, Randall Balestriero
arXiv AI
3d ago

RISED: RubrIcs for agentic multi-environment Selection and sElf-Distillation

RISED introduces a framework that uses rubric-based textual feedback to improve training of a single large language model (LLM) agent across multiple interactive environments. By having an LLM judge tag rollouts with a shared rubric vocabulary, the system guides both online data selection and policy supervision, enabling richer cross‑environment relationships and within‑group reward contrast. Experiments show that RISED achieves the highest mean pass rate and ranks first or second in every individual environment, with rubric analysis revealing behavioural changes behind these gains.

By Jingtan Wang, Sirajul Salekin, Young mok Jung, Javier Movellan, Bryan Kian Hsiang Low, Manjot Bilkhu
arXiv AI
3d ago

ShatterQuant: Breaking Uniform Precision with Block-Wise Mixed-Precision on a Systolic Transformer Hardware Accelerator

ShatterQuant is a hardware-software co-designed framework that enables mixed-precision quantization within individual tensors by assigning different bit-widths to blocks of a weight projection. It couples precision granularity with processing element configuration, allowing each precision to determine an effective block height. The framework includes a hardware-aware post-training method based on block-level standard deviation and weight sensitivity, a ShatterQuant Transformer Accelerator supporting 1/2/4/8-bit weight precision, precision-dependent PE configuration, block rescaling, and integrated softmax and piecewise-linear nonlinearities, and an evaluation showing 1.5 TOPS, 760 GOPS/$mm^2$ area efficiency, and 2.8 TOPS/W energy efficiency on a TSMC 16nm PDK implementation.

By Mikolaj Walczak, Edward Humes, Chao Fang, Marian Verhelst, Tinoosh Mohsenin