Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,731 stories · RSS feed

arXiv Machine Learning
Sep 30

Relative Kinetic Utility: Calibrating Cross-Layer Credit for Global Structured LLM Pruning

The paper introduces Global Relative Kinetic Utility (Global RKU), a label‑free method for calibrating cross‑layer credit in global structured pruning of large language models. Global RKU estimates channel importance via a final‑hidden‑state activation‑gradient signal and applies block‑relative normalization to remove block‑common scale while preserving within‑block ordering, enabling a single‑stage static pruning topology. Experiments on Qwen‑2.5‑7B show significant performance gains at various sparsity levels, and ablation studies confirm the effectiveness of the relative‑normalization step.

By Tianhao Qian, Guilin Qi, Jiayu Chen
arXiv Machine Learning
Sep 30

Alignment-Guided Flow Transformer for Efficient Vision-Language-Action Policy Learning

The paper introduces Alignment‑Guided Flow Transformer (AGFT), a framework for Vision‑Language‑Action (VLA) models that explicitly enforces tri‑modal alignment among vision, language, and action through a dedicated alignment loss. AGFT bridges representational gaps across modalities, improving task adaptation and robustness, and employs a flow‑matching objective to reduce inference steps compared to diffusion‑based policies. Experiments on a large benchmark demonstrate that AGFT achieves higher success rates and lower inference latency than state‑of‑the‑art baselines, highlighting tri‑modal alignment as crucial for scalable VLA manipulation.

By Shengchao Hu, Peng Wang, Qiyang Zhou, Guodong Zheng, Yuqi Huang, Li Shen, Ya Zhang, Dacheng Tao
arXiv Computer Vision
Sep 30

EGSD: Event-Grounded Self-Distillation for Streaming Video Understanding

The paper introduces Event‑Grounded Self‑Distillation (EGSD), a method for real‑time video understanding that treats streaming memory as incremental updates over verifiable events such as key visual entities, actions, and details. EGSD adapts on‑policy self‑distillation by combining a multiplicative weight with outcome reward to align teacher and student preferences, and re‑weights the teacher with event information plus an entity‑coverage reward to counter question‑relevance bias. Experiments on StreamingBench and OVO‑Bench Real‑Time track show EGSD achieves 79.8 % and 73.4 % accuracy respectively, while improving effective‑entity recall by 17.4 % with only a 6.8 % increase in memory length.

By Yuwei Miao, Xuesheng Zhang, Wenhao Zou, Jixia Zhang, Jianwei Lv, Bo Yuan, Junfeng Wang, Shiao Xie
arXiv Computer Vision
Sep 30

Waypoint-1.5: A Real-Time Video World Model for Consumer Hardware

Waypoint‑1.5 is a real‑time diffusion world model designed for interactive video generation on consumer‑grade hardware. It is pre‑trained on 100,000 hours of control‑aligned video game footage and can generate playable video conditioned on full keyboard and mouse input. The system offers two resolution variants, distinguishes rendered FPS, latent FPS, and control rate, and includes a detailed data pipeline, architecture, training methodology, and runtime system.

By Rajit Rajpal, Shahbuland Matiana, Liew Wei Pyn, Anmol Agarwal, Ryan Craig, Andrew Lapp, Mithun Hunsur, Sami BuGhanem, Scottie Fox, Aaron Sanders Carson Poole, Irene Park, Dave Rossi, Spencer Frazier, Louis Castricato