ComputerSD is an online self‑distillation method for computer‑use agents that leverages real‑time feedback from executed GUI transitions. It uses a fine‑tuned GUI analyzer to generate guidance and a step‑level value score after each action, combining token‑level OPSD with trajectory‑level GRPO in an asynchronous training framework. On the OSWorld‑Verified benchmark, ComputerSD improves performance over outcome‑only GRPO by 1.9 and 4.1 percentage points on Qwen3‑VL‑8B‑Thinking and EvoCUA‑8B backbones, and shows strong generalizability in out‑of‑distribution tests.
By Yong Du, Tongbo Chen, Zhengxi Lu, Yizhou Liu, Bofan Chen, Tao Jiang, Wenhao Xu, Yongliang Shen
The paper develops a statistical theory for minimum‑norm interpolation in high‑dimensional regression, showing how regularization geometry and signal sparsity affect generalization. It identifies regimes where sparsity‑promoting regularizers yield exact interpolation that is far more accurate than approximate fitting, and proves a zero–one generalization law for strongly overparameterized noiseless problems. The authors also characterize training and generalization errors along ρ‑regularization paths when feature dimension and sample size are proportional, demonstrating that generalization improves with more sparsity‑promoting norms and sparser targets, and that small changes in regularization strength can cause large shifts in generalization.
whyItMatters":"The work provides a quantitative understanding of delayed generalization (grokking) and reveals a statistical instability in minimum‑norm interpolation, offering insights that could guide the design of regularizers for better generalization in overparameterized models."
By Gil Kur, Ileana Rugina, Cl\'ementine Carla Juliette Domin\'e, Marco Mondelli
The paper investigates on‑policy distillation (OPD) versus supervised fine‑tuning (SFT), focusing on how students learn from multiple teachers by minimizing divergence. It shows that using forward KL divergence leads to a weighted arithmetic mixture, while reverse KL produces a normalized weighted geometric aggregate. The authors develop algorithms for both off‑policy and on‑policy settings, prove logarithmic regret bounds in tabular cases, extend the analysis to function approximation, and analyze how these aggregation targets explain OPD’s benefits and fragility.
By Qiwei Di, Xuheng Li, Kaixuan Ji, Chenggong Zhang, Heyang Zhao, Quanquan Gu
arXiv:2609.38672v1 Announce Type: cross
Abstract: Beam-search-based test-time methods provide an effective way to improve large language model (LLM) performance on long-horizon generation by pruning...
By Qijia He, Yu Huang, Yuan Cheng, Yuxin Chen, Yingbin Liang
The paper introduces Length Self-Distillation (LSD) to address the length‑scaling tax (LST) that occurs during reinforcement‑learning post‑training, where models produce unnecessarily verbose responses to already‑solved prompts. LSD routes solved prompts to an on‑policy distillation process while keeping the original RL objective for unsolved prompts, using an exponential moving average of the online policy as its teacher. Experiments show LSD matches or surpasses RL performance while reducing LST from 19.0% to –3.7% on single‑turn reasoning and from 31.4% to 13.7% on multi‑turn agentic tasks, thereby maintaining concise responses on easy queries while still enabling exploration on difficult ones.
By Xu Wan, Wenyue Xu, Shengjie Zhao, Mingyang Sun
The paper introduces GFD-OPD, a method for on‑policy distillation of diffusion models that addresses challenges when compressing large teachers into smaller students. It identifies that standard distillation fails due to distribution gaps and classifier‑free guidance amplification, and proposes Fixed‑State KL to measure these gaps. GFD‑OPD reduces the student‑teacher discrepancy and achieves state‑of‑the‑art performance across multiple benchmarks.
By Zhenxing Zhang, Jiayan Teng, Wenxu Wu, Zhuoyi Yang, Jiazheng Xu, Wendi Zheng, Jie Tang, Dan Guo, Meng Wang
arXiv:2609.40235v1 Announce Type: cross
Abstract: Continuous diffusion language models generate all tokens in parallel, yet high-quality generation can still require hundreds of network evaluations (...
By Paul Le Van Kiem, Dario Shariatian, Umut Simsekli, Alain Durmus
arXiv:2609.40316v1 Announce Type: cross
Abstract: Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed p...
By Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
arXiv:2609.38263v1 Announce Type: new
Abstract: Feature selection in neural networks remains a challenging problem, particularly in the presence of noisy or contaminated data. LassoNet is a recent ap...
By Daniela De Canditiis, Italia De Feis, Paola Stolfi
arXiv:2609.38477v1 Announce Type: cross
Abstract: Large language models (LLMs) incur substantial storage, memory-bandwidth and energy costs, motivating compact weight representations. Existing seed-b...
By Qiuyu Ren, Sudipta Paria, Aritra Dasgupta, Swarup Bhunia
arXiv:2609.38792v1 Announce Type: new
Abstract: We study training LLM judges from natural language feedback, especially for subjective tasks where the verdict depends strongly on which evaluation cri...
By Ilgee Hong, Changlong Yu, Zhenghao Xu, Xin Liu, Yuwei Zhang, Qin Lu, Bing Yin, Tuo Zhao
CellMSA introduces a novel single‑cell representation learning framework that leverages a multiple‑sequence‑alignment‑inspired context model. For each target cell, it retrieves relevant cells across batches and related cell types, summarizing cross‑cell patterns into a context‑dependent gene‑pair representation that is fed into a pair‑aware encoder. Pretraining on a massive human single‑cell corpus (≈109 million cells) and subsequent benchmarks demonstrate consistent performance gains over existing methods.
By Suyuan Zhao, Minghao Liu, Yizhen Luo, Zaiqing Nie
arXiv:2609.39229v1 Announce Type: cross
Abstract: Automatic evaluation of faithfulness increasingly relies on a large language model acting as a judge, yet the most reliable judges are proprietary fr...
By Elia Onofri, Roberto Di Pietro
The paper investigates transferring specialized task-oriented behavior to a general language model without training or distillation. It applies two training‑free heterogeneous merging techniques—Intersection‑Merge (IM) and Activate‑Prune‑Merge (APM)—to project a specialist donor into the recipient’s parameter space and interpolate backbone weights. Experiments across embedding, reranking, reward modeling, and MoE code‑specialist tasks show that both methods improve the general model, demonstrating that simple parameter‑level merging can transfer capabilities across diverse specialist roles.
By Jiahe Fan, Si Chen, Yinghao Hou, Wenbo Xia, Ke Xu, Hong Xie, Enhong Chen
arXiv:2609. 39449v1 Announce Type: new Abstract: Distributionally robust optimization (DRO) studies parameter estimation under uncertainty in the underlying probability distribution and has emerged as a principled framework for analyzing robustness and generalization.
By Elis Stefansson, David V\"avinggren, Ant\^onio H. Ribeiro
arXiv:2609.39687v1 Announce Type: new
Abstract: On-policy self-distillation (OPSD) trains mathematical reasoning models using a privileged teacher that sees a reference solution and supervises studen...
By Xincheng Wei, Yifan Ding, Yoshua Li, Yuquan Lu, Ziheng Li, Yi Lu, Dongsheng Ma, Rongxiang Weng, Xunliang Cai
arXiv:2609.39884v1 Announce Type: cross
Abstract: Role prompting elicits specialized behavior from large language models through an expert identity, offering a lightweight way to guide reasoning on d...
By Weijie Ren, Yanwen Zhang, Hao Li, Zhuolin Qi, Hengyi Zhang, Naibo Wang
MeanVoiceFlow2 is a new voice conversion framework that jointly optimizes a flow-based conversion module and a computationally efficient content encoder. It is trained via conversion distillation from MeanVoiceFlow and real data reconstruction, and further enhanced with diffusion-GAN training, sample mixing, and teacher-guided conditioning augmentation. Experiments on zero-shot voice conversion show that MeanVoiceFlow2 delivers higher perceptual quality and about nine times faster inference than its predecessor while preserving speaker similarity.
By Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Yuto Kondo
The study evaluates whether a fine‑tuned open‑weight model (Gemma‑3‑12B) can match the performance of GPT‑4o in extracting multi‑label intracranial hemorrhage acuity from non‑contrast head‑CT reports. Using a 2×2 design that varied adaptation strategy (classification head vs. instruction fine‑tuning) and training‑data source (distilled real GPT‑4o labels vs. synthetic GPT‑4o‑generated reports), the distilled instruction‑tuned model achieved macro‑F1 scores comparable to GPT‑4o and surpassed the untuned base model. The key finding is that the source of training data—distilled real reports—was more important than the fine‑tuning method, and that the entire fine‑tuning and inference process fits on a single 24 GB consumer GPU.
By Aawez Mansuri, Kush Mehta, Mohammadreza Chavoshi, Jahanzaib Malik, Theodorus Dapamede, Frank Li, Rohan Isaac, Beatrice Brown-Mulry, Chiratidzo Rudado Sanyika, YoungSeok Jeon, Judy W. Gichoya, Ali Emami, Hari Trivedi
The paper presents efficient approximations for key statistics of the Neural Tangent Kernel (NTK) in finite-width neural networks using randomized trace estimation (Hutch++). It demonstrates that the NTK trace, Frobenius norm, effective rank, and alignment can be estimated with high accuracy via matrix-free products, leveraging the NTK’s positive-semidefinite structure to use one-sided estimators with forward or reverse-mode differentiation. Experiments on MLPs, GRUs, and a 410‑million‑parameter Transformer show orders‑of‑magnitude speedups and enable practical state‑space NTK diagnostics at large scales, including applications to RNN training and data‑scarce knowledge distillation.
By James Hazelden, Balaaji Reddy Nagireddy, Eric Shea-Brown