arXiv:2609.36262v1 Announce Type: new
Abstract: Recent studies have observed that parameter changes during language-model post-training can be concentrated in a small subset of coordinates. This phen...
By Yufan Zhang, Sagnik Mukherjee, Hao Peng
arXiv:2609.36654v1 Announce Type: new
Abstract: Large language models make weight storage and memory traffic major inference costs, motivating low-precision formats that represent each weight with on...
By Ruiyi Ding, Jie Li, Kang He, Ziyan Liu, Chengru Song, Yuedong Xu, Yuan Cheng
arXiv:2609.36608v1 Announce Type: new
Abstract: On-policy distillation (OPD) trains multi-turn language agents with dense teacher supervision on student-generated responses. However, standard think-t...
By Zubin Zheng, Jiahao Wu, Shaofeng Zhang, Zhirui Zhang, Yew-Soon Ong, Shengcai Liu
arXiv:2609.37416v1 Announce Type: new
Abstract: Post-training quantization (PTQ) methods in the GPTQ family minimize a layer-wise reconstruction error on a uniform grid whose scale must be chosen; th...
By Jonas von Berg, Massimiliano Datres, Carlo Knei{\ss}l, Gitta Kutyniok
arXiv:2609.37522v1 Announce Type: new
Abstract: On-policy distillation (OPD) trains compact language agents with teacher feedback on student-generated trajectories. In multi-turn tasks, compounding e...
By Xiaohan Yi, Wen Luo, Yani Huang, Junfeng Zhan, Asher Qin, Peilin Zhao, Xi Xiao
arXiv:2609.37887v1 Announce Type: new
Abstract: Activation and key-value cache precision change what a quantized language model computes without altering its stored weights. Direct weight-code bounds...
By Arian Eamaz, Mojtaba Soltanalian
arXiv:2609.37924v1 Announce Type: new
Abstract: Recent work on anchored diffusion language models improves denoising by shaping an intermediate latent space with supervised important-token targets. I...
By Joel Anto Paul, Litu Rout, Aditya Akella, Sanjay Shakkottai
arXiv:2609.36540v1 Announce Type: cross
Abstract: Generalist robot policies such as vision-language-action models (VLAs) have achieved remarkable generalization, but their inference delays can confli...
By Moritz Zoellner, Reece O'Mahoney, Ioannis Havoutis, Rohan Paleja
arXiv:2609.37202v1 Announce Type: cross
Abstract: Droplet collision governs droplet population dynamics in many chemical engineering processes, such as spray drying, spray cooling, agricultural spray...
By Weiming Xu, Tao Yang, Peng Zhang
arXiv:2609.37447v1 Announce Type: cross
Abstract: How strong can an AlphaZero-style chess system become under limited training compute when its entire learning loop is engineered for efficiency? We t...
By Bertil Braun
arXiv:2609.37532v1 Announce Type: cross
Abstract: Growing large language model applications demand efficient inference. At high concurrency, block-diffusion speculative decoding suffers from verifica...
By Rongjian Chen, Minxian Xu, Zhengxin Fang, Kejiang Ye, Chengzhong Xu
arXiv:2609.37581v1 Announce Type: cross
Abstract: Vision-Language Models (VLMs) excel at visual understanding and reasoning but often incur substantial inference costs due to the large number of visu...
By Jing Wang, Zhiping Wu, Dongdong Ren, Youfang Han, Wei Zhao, Wenbin Li
arXiv:2609.36734v1 Announce Type: new
Abstract: Knowledge Distillation (KD) trains a smaller-capacity student model to imitate a larger-capacity teacher model by matching output distributions, implic...
By Ayan Sengupta, Vaibhav Seth, Tanmoy Chakraborty
arXiv:2609.36965v1 Announce Type: new
Abstract: System One models such as Jev offer an efficient alternative to generative language models for tasks that require decisions rather than open-ended resp...
By Zexiao Wang, Zihao Zhang, Xudong Wang, Pan Wang, Ziyi Ye, Haoyu Zhao, Zuxuan Wu, Shuicheng Yan
arXiv:2609.36374v1 Announce Type: new
Abstract: Feed-forward 3D reconstruction models have achieved impressive performance by scaling model and dataset size, but their cost excludes most research gro...
By Brandon Leblanc, Charalambos Poullis
arXiv:2609.37481v1 Announce Type: new
Abstract: Large-scale pervasive sensing increasingly relies on high-resolution satellite imagery, yet task-specific onboard vision is constrained by costly annot...
By Ahmed Abdelnaby, Mohamed Elmahallawy, Marius Bernahrndt, Tobias Hecking
arXiv:2604.08995v3 Announce Type: replace
Abstract: With the advancement of interactive video generation, diffusion models have increasingly demonstrated their potential as world models. However, exi...
By Zile Wang, Zexiang Liu, Jiaxing Li, Kaichen Huang, Baixin Xu, Fei Kang, Mengyin An, Peiyu Wang, Biao Jiang, Yichen Wei, Yidan Xietian, Jiangbo Pei, Liang Hu, Boyi Jiang, Hua Xue, Zidong Wang, Haofeng Sun, Wei Li, Wanli Ouyang, Xianglong He, Yang Liu, Yangguang Li, Yahui Zhou
The paper investigates how the quantity, source, and selection of prompts influence transfer in on‑policy distillation (OPD) between teacher and student models. It shows that a small set of well‑chosen prompts can achieve performance comparable to large prompt pools, but the effectiveness of prompts depends on the specific teacher‑student pair and target task. The study also finds that prompt utility is relational rather than intrinsic, and that targeted prompt selection does not consistently outperform random sampling.
By Jiaxuan Wang, Jiafei Lyu, Yuchen Cai, Siye Wu, Pengyuan Wang, Jiashun Liu, Xiang Cheng, Kai Yang, Yangkun Chen, Saiyong Yang, Lan-Zhe Guo
arXiv:2609.37817v1 Announce Type: new
Abstract: Geometric representation learning predominantly scaffolds representations onto flat Euclidean subspaces or compact product tori ($\mathbb{T}^K$). Howev...
By Zhongping Ji
arXiv:2609.37969v1 Announce Type: new
Abstract: High-resolution video generation is expensive, as its cost grows rapidly with the number of spatiotemporal tokens. A practical alternative first genera...
By Haozhe Liu, Tian Ye, Shuchen Xue, Yitong Li, Junsong Chen, Haopeng Li, Jincheng Yu, Duomin Wang, Ruihua Zhang, Lei Zhu, Song Han, Enze Xie