arXiv:2609.37132v1 Announce Type: new
Abstract: On-policy self-distillation (OPSD) improves large language models by letting a self-teacher with privileged information provide dense token-level super...
By Zheng Zhang, Xinyue Tan, Lufei Li, Xinyi Zhang, Yexin Li, Kan Ren
arXiv:2609.37670v1 Announce Type: new
Abstract: MeanFlow enables efficient few-step generation by predicting interval-average velocities, but this representation creates a mismatch for reward fine-tu...
By Haocheng Tang, Tianchi Xie, Xingqiao Lin
arXiv:2609.37673v1 Announce Type: new
Abstract: Experienced professionals know more than just facts and conclusions. They know which cues matter, why a judgment is reasonable, and which action to tak...
By Changmian Wang, Yuchao Ma, Xuchao Lu, Chen Zhang, Ping Sun, Jiazheng Wang, Shan Wang, Xuanwen Chen, Yihe Sun, Ziyu Lu, Jianqiang Huang, Hongzhi Li, Ziqing Xia, Kaihua Tang, Xian-Sheng Hua, Qinghua Zheng
arXiv:2609.37829v1 Announce Type: new
Abstract: Video diffusion transformers (DiTs) increasingly adopt mixture-of-experts (MoE) architectures to reduce active computation, but their full expert stora...
By Jiachang Zhang, Teng Hu, Bohao Feng, Songhang Shen, Wenqiang Wang, Hongqian Deng, Ran Yi
arXiv:2609.37898v1 Announce Type: new
Abstract: Reinforcement learning for long-horizon agents typically relies on sparse outcome-based rewards. This leads to a severe cold-start problem, as early-st...
By Youling Huang, Tiankuo Xu, Jiaji Liu, Tong Zheng, Shuo Zhou, Shaotong Qi, Junchi Yao, Shiyang Liu, Hao Xu, Pengcheng Xu, Bo Huang, Hongyi Fu, Lin Lin
arXiv:2609.38142v1 Announce Type: new
Abstract: A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advi...
By Rishabh Agrawal, Hejie Cui, Shasha Li, Shanchan Wu, Sercan \"{O}. Ar{\i}k
arXiv:2609.36049v1 Announce Type: cross
Abstract: Worker-monitor setups are a promising approach to AI oversight, but training workers against fixed monitors can incentivize monitor evasion. We study...
By Joseph H. Rudoler, Kevin Tan, Benedict Tessler, Timothy Kong, Enric Boix Adser\`a
arXiv:2609.36173v1 Announce Type: cross
Abstract: Parallel speculative drafting generates multiple candidates in one backbone pass, but independent token selection can produce inconsistent continuati...
By Haohui Zhang, Keyu Chen, Haocheng Sun, Weibo Gu, Ruizhi Qiao, Xing Sun, Bo Jiang
arXiv:2609.36376v1 Announce Type: cross
Abstract: Dense retrieval, the key component of Retrieval Augmented Generation (RAG), retrieves the most relevant documents by comparing dense vector represent...
By Louis Tremblay Thibault, Sofiane Azogagh, Marc-Olivier Killijian, Ulrich A\"ivodji
arXiv:2609.36471v1 Announce Type: cross
Abstract: World-Action Models (WAMs) improve robotic manipulation by conditioning action generation on predicted future observations, but future prediction add...
By Guoheng Sun, Chen Chen, Jin Wang, Ang Li, Teresa Lv
arXiv:2609.36546v1 Announce Type: cross
Abstract: On-policy distillation (OPD) trains a student model on its self-generated trajectories with dense token-level teacher feedback. However, naive OPD ma...
By Shutong Wu, Xiwen Chen, Brendan Rappazzo, Daiheng Zhang, Anderson Schneider, Yuriy Nevmyvaka, Jiawei Zhang
arXiv:2609.36879v1 Announce Type: cross
Abstract: As LLM-based agents perform increasingly complex tasks, Agent Skills have emerged as a flexible mechanism for extending their capabilities. An Agent...
By Haoran Ou, Gelei Deng, Xuanye Zhang, Wenbo Guo, Tianwei Zhang, Kwok-Yan Lam
arXiv:2609.36954v1 Announce Type: cross
Abstract: Distributed inference depends on GPU collective communication that must keep pace with evolving hardware and specialized workloads. However, existing...
By Osayamen Jonathan Aimuyo, Swapnil Gandhi, Christos Kozyrakis
arXiv:2609.37041v1 Announce Type: cross
Abstract: Self-Distillation Fine-Tuning (SDFT) enables a language model to act as its own teacher: by conditioning on a demonstration, the model produces an im...
By Su Ee Tan, Xiaotong Ji, Rasul Tutunov, Haitham Bou-Ammar, Matthieu Zimmer
arXiv:2609.37702v1 Announce Type: cross
Abstract: Class Incremental Learning (Class-IL) requires models to learn new classes over time while preserving previously acquired knowledge without access to...
By A. L. S. Conde, Y. Elkhatib, C. M. Ranieri
arXiv:2609.37916v1 Announce Type: cross
Abstract: Production machine learning (ML) stacks often split graph compilation and kernel execution across different layers and languages, making backend beha...
By Eugene Hauptmann, Nataliya Kosmyna
arXiv:2609.34768v2 Announce Type: replace
Abstract: Millimeter-wave (mmWave) radar enables privacy-preserving human perception, but the extreme sparsity of point clouds from commercial single-chip se...
By Shuxing Zhang, Yongquan Ni, Zhenyu Ding, Yawen Lin
arXiv:2605.06140v3 Announce Type: replace-cross
Abstract: Generative modeling of physical systems, such as molecules, requires learning distributions that are invariant under global symmetries, such...
By Samir Darouich, Vinh Tong, Llu\'is Pastor-P\'erez, Tanja Bien, Loay Mualem, Mathias Niepert
arXiv:2607.26621v3 Announce Type: replace-cross
Abstract: Large language models (LLMs) have demonstrated strong reasoning capabilities, motivating their use as the backbone of foundation recommendati...
By Hao Jiang, Peiru Du, Pengfei Yao, Mengting Li, Siyuan Lou, Kuo Cai, Sheng Yu, Qiang Luo, Jian Liang, Ruiming Tang, Fei Pan, Peng Jiang, Wenwu Ou
arXiv:2609.35491v2 Announce Type: replace-cross
Abstract: Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion tea...
By Chi Zhang, Yueyi Liu, Haoyang Shi, Ruichuan An, Haoyu Li, Yuhang Wu, Sen Cui, Miao Liu