Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,834 stories · RSS feed

arXiv Computation and Language
Sep 23

Prompt Breadth and Rollout Refresh Interact in On-Policy Distillation

The study investigates how the number of prompts and the strategy of refreshing rollout responses affect on‑policy distillation (OPD). Using a 3×3 experiment with 14,080 trajectories and 110 optimizer updates, the authors find that with ten policy snapshots, eight prompts achieve 24.09% accuracy—nearly matching the 24.51% obtained with 14,080 distinct prompts. However, when responses are frozen at the initial policy, increasing prompt breadth actually reduces accuracy, whereas per‑update refresh raises it, producing a 4.07‑point interaction effect. Comparisons with two teacher models show that periodic models excel in short‑budget accuracy and answer completion, but frozen‑response models surpass them in overall accuracy at a 32K output limit, using 1.7–1.8× more response tokens. whyItMatters":"The findings demonstrate that prompt efficiency in OPD is contingent on both the refresh strategy and the inference budget, informing how to design more effective distillation pipelines."

By Lingxiang Hu, Tianle Xia, Ming Xu, Yiding Sun, Linfang Shang
arXiv Computation and Language
Sep 23

LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay

The paper introduces LatentPort, a method that allows a language model to transfer its live memory to another model without requiring the receiver to reread the context. Experiments on a Qwen3.5 4B-to-9B sibling pair show that adding a Gated DeltaNet (GDN) persistent-state package reduces negative log‑likelihood by 0.747 nats/token and improves performance across 64 PG19 documents. The study also demonstrates that direct recurrent and convolution reuse outperforms learned GDN maps, and a 434,176‑parameter correction further narrows the performance gap to the native 9B model.

By Simon P. Villani
arXiv Computation and Language
Sep 23

Syndrome, Synergy, and Safety: Structured Reasoning and Knowledge-Driven Alignment for TCM Prescription Generation

The paper introduces a four‑stage framework—SFT, PG‑CoT, Dynamic, and K‑RL—to improve large language models for Traditional Chinese Medicine prescription generation. It addresses three key gaps: lack of auditable reasoning (SR Gap), failure to adjust prescriptions over time (LA Gap), and non‑enforcement of absolute contraindication rules (SC Gap). Experiments on 12 fine‑tuned models and 6 zero‑shot baselines show that the framework, particularly a 7B Mistral model, outperforms zero‑shot GPT‑5 on all three TCM evaluation metrics.

By Zheng Chen, ZhiCheng Du, Haoxuan Li, Peiwu Qin
arXiv Computation and Language
Sep 23

TSS: Target-Side Sparsification for Speculative Decoding in Domain-Specific Large Language Models

The paper introduces TSS, a target-side sparsification framework that selectively skips layers in a target verifier during speculative decoding for domain-specific large language models. By exploring multi-layer skip configurations with an acceptance- and metric-aware breadth search, TSS reduces verification cost, increases draft acceptance, and can even improve downstream task performance without retraining. Experiments on Spec-Bench demonstrate consistent gains across domains and model scales, notably boosting translation throughput by 1.68× and improving BLEU scores significantly.

By Haibo Hu, Lianming Huang, Qiao Li, Nan Guan, Chun Jason Xue
arXiv Computation and Language
Sep 23

PACE-dLLM: Elastic Block Decoding via Confidence Cliff Estimation for Diffusion Language Models

The paper introduces PACE-dLLM, an acceleration method for diffusion language models (dLLMs) that uses the model’s own per‑step confidence to estimate a ‘confidence cliff’ and determine the optimal look‑ahead horizon for block decoding. By fitting this cliff in closed form at each step, PACE-dLLM sets the horizon to its saturation point and applies an independent confidence threshold for token commitment, thereby avoiding the trade‑offs inherent in fixed‑size block decoding. Experiments on reasoning and code benchmarks show that PACE-dLLM achieves the best average accuracy on open‑source dLLM backbones while delivering significant wall‑clock speedups—up to 5.23× on LLaDA and 3.06× on Dream—improving the quality‑throughput Pareto frontier.

By Xiaocheng Lu, Shuhan Guo, Ziyue Ma, Jie Zhang, Jian Liu, Jingcai Guo, Haoxuan Che, Song Guo
arXiv Computation and Language
Sep 23

Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs

Flash-dLLM is a training‑free inference acceleration framework that improves the speed and memory efficiency of Diffusion Large Language Models (dLLMs). It tackles GPU memory I/O bottlenecks by introducing an I/O‑aware fused KV‑cache kernel and then employs a draft‑and‑verify decoding strategy that uses the dLLM itself as both drafter and verifier. Experiments on mathematical reasoning and code‑generation tasks show Flash‑dLLM outperforms existing acceleration methods, achieving up to 11.0× speedups over the Elastic‑Cache baseline.

By Quan Nguyen-Tri, Mukul Ranjan, Zhiqiang Shen
arXiv Computer Vision
Sep 23

MirrorDistill: Illumination-Aware Latent Distillation for Efficient Low-Light Restoration

MirrorDistill introduces an illumination‑aware latent distillation framework for low‑light image enhancement. It trains a lightweight student encoder‑decoder by aligning its intermediate features with clean‑domain targets generated by a teacher decoder, using feature mirroring and illumination‑aware weighting to emphasize underexposed regions. The method achieves state‑of‑the‑art performance on the LOL‑v2‑Real benchmark while maintaining the lowest computational complexity, and the code is released as open source.

By Farida Mohsen, Tala Zaim, Nurul Izni Rusli, Ali Al-Zawqari, Ali Safa, Samir Brahim Belhaouari
arXiv Computer Vision
Sep 23

RGSQ: Riemannian Geometry-Sensitive Quantization for Large Vision-Language Models

RGSQ introduces a Riemannian geometry‑aware post‑training quantization method for large vision‑language models, treating quantization as a reconstruction problem under a Fisher‑Riemannian metric. It identifies modality‑specific sensitive directions via manifold mappings and applies geometry‑aligned rotations and whitening to steer low‑bit perturbations toward loss‑insensitive axes. Experiments on diverse VLM benchmarks show RGSQ delivers the best accuracy and stability in extremely low‑bit settings, outperforming existing VLM‑aware baselines by up to 5.9% and single‑modality methods by up to 8.6%.

By Zhiping Wu, Dongdong Ren, Yangchengyu Zhou, Zhengjie Zhang, Wenbin Li, Hongbing Pan, Yang Gao
arXiv Computer Vision
Sep 23

Shallow to Deep: Aligning Token Pruning with Stage-wise Roles in LVLMs

The paper introduces STD, a hierarchical token pruning framework for Large Vision‑Language Models that aligns pruning strategies with the functional roles of different network stages. By using high‑frequency spectral analysis in shallow layers, Gaussian‑smoothed attention in intermediate layers, and a stability‑adaptive trigger in deep layers, STD preserves essential visual information while aggressively reducing token counts. Experiments demonstrate that STD outperforms existing pruning methods, achieving up to 94.4% token reduction and a 3.9× speed‑up on LLaVA‑NeXT‑7B.

By Shuo Zhang, Jintao Tong, Yixiong Zou, Yuhua Li, Ruixuan Li
arXiv Computer Vision
Sep 23

Delving into Asymmetric Information Dynamics for High-Fidelity Virtual Try-On

The paper introduces RealFit, a virtual try‑on framework that addresses texture degradation and structural drift caused by symmetric bidirectional attention in Diffusion Transformers. By employing Unidirectional Information Flow to isolate garment conditions from stochastic noise and Decoupled Timestep Modulation to strengthen the conditional signal, RealFit achieves high‑fidelity garment rendering. The method also enables a time‑invariant condition branch and a conditional KV cache, cutting inference time by about 75%.

By Zishu Qin, Zhiyu Jin, Pipei Huang, Hao Zhou
arXiv Computer Vision
Sep 23

Latent Dataset Distillation for Human Motion Prediction

The paper introduces a latent dataset distillation framework for human motion prediction, addressing the limitations of traditional gradient matching by incorporating a learned motion prior. Motions are compressed using a residual‑quantized variational autoencoder, and distillation updates only a latent bank while keeping the decoder frozen, ensuring synthetic motions remain plausible. Experiments on Human3.6M, CMU, and 3DPW datasets demonstrate that this method outperforms direct gradient matching in most settings and yields more realistic synthetic motions.

By Ge Tian, Guang Li, Takahiro Ogawa, Miki Haseyama
arXiv Computer Vision
Sep 23

From Token Importance to Conditional Removability: Rethinking Visual Token Pruning in Multimodal Large Language Models

The paper argues that token importance alone is insufficient to determine safe removal of visual tokens in multimodal large language models, because removability depends on representation depth and the surrounding deletion set. Through controlled experiments, the authors show that the same tokens can have different effects when removed at different depths or contexts. They introduce CoRePrune, a training‑free two‑stage pruning framework that refreshes deletion effects as visual representations evolve and refines candidate tokens based on the current deletion set, achieving high performance retention across multiple backbones and reducing prefill time significantly.

By Shengli He, Yongchao Liang, Roumeng He, Junjie Zeng, Jiyuan He, Xin Fang, Can Wu, Li Zheng
arXiv Computer Vision
Sep 23

DIFTA-3D: Depth-Consistent Instance-Level Feature Transfer and Adaptation of DINOv3 for 3D Detection

The paper introduces DIFTA-3D, a method that replaces the task‑specific visual branch in IIFNet3D with a frozen DINOv3 foundation model for RGB‑D 3D instance detection. It employs a depth‑consistent feature pipeline that projects points into calibrated RGB‑D frames, filters features with a metric depth‑residual check, caches accepted DINOv3 features, and aggregates them within proposal‑aligned RoI grids. Extensive experiments on ScanNetV2 show that the DINOv3 control achieves mAP scores of 76.15/60.93 at IoU thresholds 0.25/0.50, while the Conservative VAID recipe improves these to 76.59/62.16, indicating a modest gain from the proposed transfer recipe.

By Linman Wang, ZiFei Zhang, Chunran Zheng, Xiwang Dong, Jiarong Lin
arXiv Computer Vision
Sep 23

GAD-MambaUNet: Direction-Group Mamba with Gradient-Adaptive DINOv3 Distillation for Lightweight Medical Image Segmentation

The paper introduces GAD-MambaUNet, a lightweight medical image segmentation network that integrates efficient local modeling, Direction-Group Graph Selective Scan (DG‑GSS) for structured information exchange, and training‑time supervision from a frozen DINOv3 teacher with Gradient‑Adaptive Distillation. GAD‑MambaUNet demonstrates a strong accuracy‑efficiency trade‑off compared to other lightweight and general segmentation methods, and ablation studies confirm the benefits of DG‑GSS and DINOv3‑GAD supervision. Future work aims to refine teacher‑student alignment and apply the framework to multi‑class and multi‑modal medical segmentation tasks.

By Fang Wang, Huitao Li, Wenhan Chao, Zheng Zhuo, Xinxin Yang