Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,731 stories · RSS feed

Hugging Face Trending Papers
Sep 28

EOPSA: Efficient On-Policy Self-Distilled Safety Alignment

EOPSA (Efficient On-Policy Self-Distilled Safety Alignment) addresses inefficiencies in On-Policy Self-Distillation (OPSD) for safety alignment by focusing training on safety-critical tokens. It introduces Adaptive Rollout Scheduling, which limits generation length based on a Teacher Rescue Rate metric, and Selective Distillation, which filters out safety-neutral tokens to concentrate gradient updates on safety-pivotal transitions. Experiments on models up to 32B parameters show that EOPSA reduces rollout computation by about 50% and backpropagates through only roughly 2% of tokens, outperforming full-token distillation baselines in safety compliance and reasoning retention.

Hugging Face Trending Papers
Sep 28

When Less Data Favors Smaller Teachers: Rethinking Teacher Capacity and Data Selection for Knowledge Distillation

The paper investigates how the optimal teacher model size for knowledge distillation changes with the amount of training data. It finds that when data is scarce, smaller teachers can outperform larger ones, a phenomenon driven by both score geometry and relational ordering of class predictions. The authors propose DVA, a data‑selection method that uses a small teacher to filter samples by difficulty and maximize diverse relational signals, achieving competitive results without relying on training dynamics.

Hugging Face Trending Papers
Sep 28

DPS: Dual-Mode Precision LLM Serving with Semi-Unified Memory

DPS (Dual-Mode Precision LLM Serving) is a system that treats model‑weight memory as elastic by using a multi‑precision representation. Under normal load it serves the full‑accuracy model, but when KV‑cache pressure spikes it switches to a lower‑precision variant and reallocates unused weight memory for KV cache blocks. Built on Semi‑Unified Memory and implemented on top of vLLM, DPS boosts sustained throughput by 2.1–3.3× and effective pass@1 by up to +41 pp over static FP16 while maintaining FP16‑class accuracy.

arXiv Computation and Language
Sep 28

Understanding the Role of Prompt Template in Knowledge Distillation for Safety Alignment

The study investigates how prompt template choices during Knowledge Distillation (KD) affect safety alignment in language models. It finds that using chat templates during KD degrades safety alignment, making models more compliant with harmful queries, while non-chat templates better preserve the base model’s internal representations. These effects are observed across LLaMA, Gemma, and Qwen families on multiple safety benchmarks.

By Anjila Budathoki, Manish Dhakal, Benjamin M. Ampel, Yi Ding
arXiv Computer Vision
Sep 28

The Shape of Events: Edge-Based Inductive Biases via Cross-Domain Distillation

The paper investigates how knowledge distillation from event cameras to RGB images can alter the inductive biases of convolutional neural networks. By transferring learning from the event domain, the authors find that models gain color invariance, a shape bias, and improved robustness to high‑frequency noise, largely due to reduced reliance on texture and increased emphasis on edge‑based object shape. These changes are evidenced by early‑layer processing differences and a spectral trade‑off between robustness to missing high‑frequency content and vulnerability to its contamination or geometric disruption.

By Soshun Kihara, Shunsuke Yasuki, Masato Taki
arXiv AI
Sep 28

Self-Play Search Distillation for Large Language Model Reasoning

Self-Play Search Distillation (SPSD) is a framework that generates superhuman synthetic data by having MuZero-like networks play board games in executable environments. The search records are converted into structured reasoning chains that serve as environment‑grounded supervision for training large language models. When applied to Qwen3‑4B‑Base, SPSD improves performance on six mathematics benchmarks from 24.1 to 36.6 and raises the win rate on unseen games from 15% to 45%.

By Lorenzo Molfetta, Wai-Chung Kwan, Giacomo Frisoni, Luca Ragazzi, Gianluca Moro, Pavlos Vougiouklis, Jeff Z. Pan, Pasquale Minervini
arXiv AI
Sep 28

Prompt-Based Continual Compositional Zero-Shot Learning

The paper introduces PromptCCZSL, a framework that enables vision‑language models to continually learn new attributes, objects, and their unique compositions while avoiding forgetting. It uses a frozen VLM backbone with prompt‑based techniques, recency‑weighted multi‑teacher distillation, and several loss functions (CAL, OPL, IDL) to maintain prior knowledge and promote diverse, distinct embeddings. Experiments on UT‑Zappos and C‑GQA show significant performance gains over existing VLM‑based and non‑VLM baselines, establishing a new benchmark for continual compositional zero‑shot learning.

By Sauda Maryam, Sara Nadeem, Faisal Qureshi, Mohsen Ali
arXiv Computer Vision
Sep 28

Structured Reasoning Agentic Framework for Interpretable Critical View of Safety Assessment

The paper introduces ReasonCVS, a structured reasoning framework for assessing the Critical View of Safety in laparoscopic cholecystectomy. It uses a Vision‑Language Model to build an Anatomical Scene Graph Abstraction and a Large Language Model–based Rationale‑Aware Reasoning Agent to verify sub‑criteria, producing a final verdict with traceable clinical rationale. Experiments on the Endoscapes‑CVS201 benchmark show ReasonCVS outperforms existing methods with a 68.1% mAP while offering interpretable, criterion‑level explanations.

By Qing Xu, Yuxiang Luo, Zhen Chen
arXiv Computer Vision
Sep 28

InternW0-$\Delta$: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data

InternW0-Δ is a unified World Action Model that integrates pretrained visual dynamics, scene semantics, 4D geometry, and motion priors within a Mixture-of-Transformers framework to generate robot actions. It leverages a frozen VLM for semantic guidance, a 4D foundation model for geometric priors, and introduces Causal Imprint to learn future-relevant scene changes without future-video rollout. The model is pretrained on a newly curated 20K‑hour heterogeneous corpus of robot and human demonstrations, achieving superior performance on simulation benchmarks and real‑robot platforms.

By Xingyu Miao, Zizun Li, Baole Fang, Kaiwen Song, Tenghui Wang, Hanxue Zhang, Yating Wang, Xudong Li, Yuping He, Xueyuan Wei, Chao Gao, Xijie Yang, Yingxiang Xu, Kerui Ren, Wenqi Guo, Jianjun Zhou, Xinzhe Wang, Weiguang Zhao, Ni Yang, Zetao Cai, Yufei Xue, Hengjie Li, Zeyu He, Yuanzhen Zhou, Rong Fu, Jianyang Zhang, Siwei Cui, Fuxian Huang, Yunsong Zhou, Xing Gao, Yifei Yao, Qiaojun Yu, Kailin Li, Ming Zhou, Mu Huang, Xinyue Li, Wenze Cui, Bingqi Jiang, Xueyue Zhu, Junting Dong, Haoyu Guo, Tao Lu, Mulin Yu, Bowen Zhou, Bin Zhao, Tianfan Xue, Weinan Zhang, Chunhua Shen
arXiv AI
Sep 28

MoSAR: Mixture of Semantic Attention Regimes for Learning Adaptive and Approximable Attention Geometries

MoSAR introduces a mixture of semantic attention regimes that learns an adaptive, distance‑dependent attention geometry from data, rather than predefining sparse or local patterns. The model uses input‑conditioned routers to select short, medium, or global regimes, creating a continuous attention field that can be discretized for efficient inference. Experiments show that MoSAR achieves lower‑reach attention without sacrificing language‑modeling quality, improving perplexity over dense RoPE and outperforming baselines like ALiBi, while remaining stable under top‑1 discretization.

By Michele Paolicelli, Alessandro Petruzzelli, Alessandro Franceso Maria Martina, Cataldo Musto, Giovanni Semeraro
arXiv AI
Sep 28

ActKV: Efficient LLM Agents through Action-Guided KV Cache Management

ActKV is a new KV cache compression framework designed for agentic large language model (LLM) inference. It prioritizes cache entries that contribute to action generation, using action-oriented eviction, confidence-driven budget allocation, and page-aware compression to reduce memory usage while preserving accuracy. In long-trace tasks, ActKV retains 98.53% of FullKV’s accuracy using only 25.98% of its peak memory and boosts token and task throughput by 3.97× and 3.58×, respectively.

By Zihan Wang, Cheng Tang, Lei Gong, Chao Wang, Wenqi Lou, Teng Wang, Xuehai Zhou
arXiv Computation and Language
Sep 28

Recursive Self-Improvement via On-Policy Distillation for Reasoning

The paper introduces a recursive self-improvement framework for language models that replaces an external teacher with a frozen copy of the student, enabling dynamic co-evolution (DCE) and self-refined concise learning (SRCL). DCE allows the privileged teacher to evolve alongside the student, while SRCL trains on shorter, verified rewrites to reduce verbosity. Experiments show that the combined DCE+SRCL approach outperforms traditional on‑policy self‑distillation across multiple model sizes and math benchmarks, achieving significant accuracy gains and shorter outputs.

By Shangjian Yin, Zehao Zhao, Kavosh Asadi, Rui Liu, Yuchen Lu, Shike Mei, Hang Cui, Luke Simon, Zhouxing Shi, Hamed Firooz
arXiv Computation and Language
Sep 28

I-Parakeet: Integer-Only Conformer ASR on Mobile NPU

I-Parakeet is an integer‑only implementation of NVIDIA’s Parakeet‑CTC Conformer ASR model that runs entirely on a smartphone NPU without any floating‑point operations or CPU fallback. The paper introduces three key techniques: an integer formulation of relative‑positional self‑attention, a minimax‑optimized Swish approximation, and layer‑wise range analysis with INT16 BatchNorm and percentile calibration for pre‑encoder activations. The resulting model achieves 4.97% WER on LibriSpeech test‑other and runs 7.5× faster than a CPU baseline on a Qualcomm NPU.

By Taichi Nishimura
arXiv Computation and Language
Sep 28

Strategically Diverse Sampling for Self-Training

The paper introduces two new sampling techniques—GROOT and Verbalized Sampling—to create strategically diverse training data for self-training large language models. By focusing on substantive variation in problem-solving approaches rather than just correctness, the authors demonstrate that models trained on this data outperform those trained on IID samples across competitive programming and Next‑Chapter Prediction tasks. Notably, self‑training with strategically diverse but incorrect traces from Qwen3‑4B surpasses IID distillation from a 235B teacher, challenging assumptions about the importance of correctness and teacher scale.

By Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata
arXiv Computer Vision
Sep 28

CSCWD: Cross-Scale Channel-wise Knowledge Distillation for Lightweight Tiny Object Detection on Edge Devices

The paper introduces Cross-Scale Channel-wise Knowledge Distillation (CSCWD), a training-time framework that transfers high‑resolution spatial representations from a YOLO11m‑P2 teacher to a lightweight YOLO11n student without changing the student’s inference architecture. CSCWD aligns teacher P2 features with student P3 while also applying same‑scale distillation at deeper pyramid levels, yielding a 2.92‑point mAP@0.5 improvement over the baseline and a 2.09‑point gain over same‑scale distillation alone. In zero‑shot tests on DUT‑Anti‑UAV and on a Raspberry Pi 5, the 2.58‑million‑parameter student reaches 50.32% mAP@0.5 at 82.32 ms latency (12.15 fps) with negligible runtime or memory increase.

By Amir Zamani, Zeinab Ghasemi-Naraghi
arXiv Computer Vision
Sep 28

PhoenixSR: Generative Heterogeneous Distillation Unleashes Efficient Models for Real-World Super-Resolution

arXiv:2609.30988v1 Announce Type: new Abstract: Real-world image super-resolution (SR) requires recovering perceptually realistic high-resolution images from complex low-resolution observations while...

By Xin Di, Mingyu Shi, Yuanfei Bao, Long Peng, Yue Zhao, Jiaming Guo, Renjing Pei, Xueyang Fu, Yang Cao, Zheng-Jun Zha
arXiv Computer Vision
Sep 28

LadderMIL: Multiple Instance Learning with Coarse-to-Fine Self-Distillation

arXiv:2502.02707v5 Announce Type: replace Abstract: Multiple Instance Learning (MIL) for whole slide image (WSI) analysis in computational pathology often neglects instance-level learning as supervis...

By Shuyang Wu, Yifu Qiu, Ines P. Nearchou, Sandrine Prost, Jonathan A. Fallowfield, Hideki Ueno, Hitoshi Tsuda, David J. Harrison, Hakan Bilen, Timothy J. Kendall
arXiv AI
Sep 28

Rank-Reliable Teacher-Guided Fitness Approximation for Expensive Evolutionary Optimization: A TinyML Architecture Search Study

The paper introduces TGL-NSGA-II, a low‑fidelity framework that uses a pretrained teacher to stratify samples by difficulty and class, then applies a short knowledge‑distillation step (KD‑Lite) before scoring candidates on a stratified evaluation set. The teacher‑guided scores are fused with a Gaussian‑process surrogate to select candidates for full evaluation, and the method is evaluated on keyword spotting and bird‑call classification tasks. Results show high Kendall‑τ values (0.74 and 0.62), a 41% reduction in proxy‑score variance, and improved hypervolume and false‑positive rates compared to full NSGA‑II, while running 2.2× faster under a constrained evaluation budget.

By Soumen Garai, Suman Samui
arXiv AI
Sep 28

CRC-Router: Risk-Constrained Routing for Medical Agentic AI Systems

CRC‑Router is a risk‑constrained, uncertainty‑aware routing module designed for medical AI systems, particularly in imaging. It fuses multiple uncertainty signals with predictive scores to estimate a per‑finding wrong‑accept risk, then uses Conformal Risk Control to set acceptance thresholds that meet a user‑specified risk target. Applied to chest X‑ray triage on the NIH ChestX‑ray14 dataset, CRC‑Router outperforms baseline methods in risk–coverage trade‑off, both as a standalone layer and when integrated with the MedRAX agent, demonstrating its effectiveness and model‑agnostic compatibility.

By Xueyang Li, Mingze Jiang, Gelei Xu, Jun Xia, Ching-Hao Chiu, Mengzhao Jia, Danny Z. Chen, Yiyu Shi