The paper introduces Global Relative Kinetic Utility (Global RKU), a label‑free method for calibrating cross‑layer credit in global structured pruning of large language models. Global RKU estimates channel importance via a final‑hidden‑state activation‑gradient signal and applies block‑relative normalization to remove block‑common scale while preserving within‑block ordering, enabling a single‑stage static pruning topology. Experiments on Qwen‑2.5‑7B show significant performance gains at various sparsity levels, and ablation studies confirm the effectiveness of the relative‑normalization step.
By Tianhao Qian, Guilin Qi, Jiayu Chen
arXiv:2605.12070v3 Announce Type: replace-cross
Abstract: Asynchronous reinforcement learning improves rollout throughput for large language model agents by decoupling sample generation from policy o...
By Zhong Guan, Yongjian Guo, Haoran Sun, Wen Huang, Shuai Di, Likang Wu, Xiong Jun Wu, Hongke Zhao
arXiv:2605.23200v2 Announce Type: replace-cross
Abstract: The linear growth of the Key-Value (KV) cache is a critical bottleneck in long-form LLM inference. Existing KV compression methods mitigate t...
By Junzhe Yang, Xiaoyu Shen
The paper introduces Alignment‑Guided Flow Transformer (AGFT), a framework for Vision‑Language‑Action (VLA) models that explicitly enforces tri‑modal alignment among vision, language, and action through a dedicated alignment loss. AGFT bridges representational gaps across modalities, improving task adaptation and robustness, and employs a flow‑matching objective to reduce inference steps compared to diffusion‑based policies. Experiments on a large benchmark demonstrate that AGFT achieves higher success rates and lower inference latency than state‑of‑the‑art baselines, highlighting tri‑modal alignment as crucial for scalable VLA manipulation.
By Shengchao Hu, Peng Wang, Qiyang Zhou, Guodong Zheng, Yuqi Huang, Li Shen, Ya Zhang, Dacheng Tao
arXiv:2501.12632v3 Announce Type: replace-cross
Abstract: Weakly supervised object localization (WSOL) models can predict both the object class and the spatial regions corresponding to the object, wi...
By Shakeeb Murtaza, Soufiane Belharbi, Alexis Guichemerre, Marco Pedersoli, Eric Granger
arXiv:2605.11608v2 Announce Type: replace-cross
Abstract: A single base LLM now comes with dozens of post-training variants, quantized, LoRA-adapted, or distilled, and each has to be checked before r...
By Chieh-Yen Lin, Shao-Hua Sun
arXiv:2605.20624v2 Announce Type: replace-cross
Abstract: Diffusion models provide powerful priors for zero-shot video inverse problems, but their real-time deployment is hindered by two inefficienci...
By Taesung Kwon, Jonghyun Park, Hyungjin Chung, Jong Chul Ye
arXiv:2609.35232v2 Announce Type: replace-cross
Abstract: Visual-token compression is effective for improving the efficiency of vision-language models, but under extreme compression budgets, token pr...
By Rui Zhong, Yu Li, Zheyu Yan, Cheng Zhuo
The paper introduces Event‑Grounded Self‑Distillation (EGSD), a method for real‑time video understanding that treats streaming memory as incremental updates over verifiable events such as key visual entities, actions, and details. EGSD adapts on‑policy self‑distillation by combining a multiplicative weight with outcome reward to align teacher and student preferences, and re‑weights the teacher with event information plus an entity‑coverage reward to counter question‑relevance bias. Experiments on StreamingBench and OVO‑Bench Real‑Time track show EGSD achieves 79.8 % and 73.4 % accuracy respectively, while improving effective‑entity recall by 17.4 % with only a 6.8 % increase in memory length.
By Yuwei Miao, Xuesheng Zhang, Wenhao Zou, Jixia Zhang, Jianwei Lv, Bo Yuan, Junfeng Wang, Shiao Xie
Waypoint‑1.5 is a real‑time diffusion world model designed for interactive video generation on consumer‑grade hardware. It is pre‑trained on 100,000 hours of control‑aligned video game footage and can generate playable video conditioned on full keyboard and mouse input. The system offers two resolution variants, distinguishes rendered FPS, latent FPS, and control rate, and includes a detailed data pipeline, architecture, training methodology, and runtime system.
By Rajit Rajpal, Shahbuland Matiana, Liew Wei Pyn, Anmol Agarwal, Ryan Craig, Andrew Lapp, Mithun Hunsur, Sami BuGhanem, Scottie Fox, Aaron Sanders Carson Poole, Irene Park, Dave Rossi, Spencer Frazier, Louis Castricato
arXiv:2609.37098v1 Announce Type: cross
Abstract: Vehicle-infrastructure cooperation can complement onboard sensing with broader and more informative observations of the traffic environment, providin...
By Junwei You, Weizhe Tang, Can Wang, Yan Zhao, Jun Hua, Haotian Shi, Wei Zhang, Lin Wang, Bin Ran
arXiv:2609.37013v1 Announce Type: cross
Abstract: Rapid assessment of building damage after natural disasters is essential to support emergency response. Earth Observation satellites can acquire rele...
By Thomas Goudemant, Benjamin Francesconi, Marjorie Bellizzi, Adrien Dorise
arXiv:2609.35882v1 Announce Type: cross
Abstract: Genome foundation models are most useful where sequences are generated, yet the largest models need datacenter accelerators and a place to send priva...
By Guang Yang, Fengchen Liu
Modern multimodal models bring generation and understanding into a single unified system, which enables them to provide and learn from their own feedback. Motivated by this unified capacity, we introd...
On-policy distillation (OPD) trains a student on its own generated responses using dense, token-level supervision from a stronger teacher. Vanilla OPD treats all teacher signals equally, assuming that...
Few-step flow-map generators, including MeanFlow and consistency models, enable efficient sampling through long-range transport, yet their on-policy distillation remains underexplored. We introduce Fl...
Multimodal on-policy distillation (OPD) aims to improve visual grounding while preserving the strong reasoning capabilities of language models. Recent multi-teacher approaches combine LLM and VLM teac...
LLM-based multi-agent systems (MAS) increasingly use latent collaboration to avoid the information loss and repeated encoding-decoding overhead of natural-language communication. However, directly for...
Chinese Semantic Error Correction (CSEC) targets semantic errors in Chinese text, which are typically more subtle and complex than spelling and grammatical errors but remain relatively underexplored....
Recursive self-improvement (RSI) aims to achieve compounding gains by having models improve themselves. While most existing RSI systems optimize external agent harnesses or prompts around a frozen bas...