arXiv Machine Learning

WinDOM: Self-Family Distillation for Small-Model GUI Grounding

arXiv:2606. 25964v1 Announce Type: cross Abstract: Small ($\sim$2B) GUI-grounding agents are attractive for on-device deployment, accessibility tooling, and low-cost iteration, but at this scale they face two open recipe questions: how to obtain bounding-box training data without expensive human annotation, and how to combine supervised fine-tuning with reinforcement learning.

arXiv AI
Sep 25

From Self-Distillation to Self-Practice: Privileged Information for Multi-Turn Agents

The paper introduces Privileged Self-Practice (PSP), a method that retains privileged information (PI) in the prompt rather than the loss during on‑policy self‑distillation for multi‑turn agents. PSP injects short per‑task instructions from an analyzer model when rollouts fail, sampling again with the instruction in context and training with the unchanged GRPO objective. Experiments on AppWorld and SWE‑bench Verified show PSP consistently outperforms plain GRPO, boosting task‑goal completion by up to 65% and resolved rate by up to 61% across three student models.

By Xingyu Su, Abhishek Kumar, Qing Ping, Youzhi Luo, Jonathan Buck, Zach Zhang, Subramanian Chidambaram, Vinayak Arannil
arXiv AI
Jul 3

DemoPSD: Disagreement-Modulated Policy Self-Distillation

arXiv:2607. 02502v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) has emerged as a practical method for training large language models (LLMs) to reason, where a single model acts as both the teacher and the student with different levels of information access.

By Yunhe Li, Hao Shi, Wenhao Liu, Mengzhe Ruan, Hanxu Hou, Zhongxiang Dai, Shuang Qiu, Linqi Song
arXiv Computation and Language
Sep 3

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients

The paper introduces Zone of Proximal Policy Optimization (ZPPO), a method that keeps a teacher model inside prompts rather than in the policy gradient to improve knowledge distillation for small students. ZPPO creates two types of reformulated prompts—Binary Candidate-included Questions (BCQ) and Negative Candidate-included Questions (NCQ)—to expose students to correct and incorrect responses, and uses a replay buffer to focus training on hard questions until the student’s accuracy improves. Experiments on the Qwen3.5 family with a 27B teacher across 31 benchmarks show that ZPPO outperforms both off‑policy and on‑policy distillation methods, especially at the smallest student scales.

By Byung-Kwan Lee, Ximing Lu, Shizhe Diao, Minki Kang, Saurav Muralidharan, Karan Sapra, Andrew Tao, Pavlo Molchanov, Yejin Choi, Yu-Chiang Frank Wang, Ryo Hachiuma
Hugging Face Trending Papers
Sep 24

From Self-Distillation to Self-Practice: Privileged Information for Multi-Turn Agents

The paper examines on‑policy self‑distillation (OPSD) for multi‑turn agents, showing that using privileged information (PI) in the loss can make agents appear confident yet underperform plain RL, sometimes worse than the untrained base model. To address this, the authors propose Privileged Self‑Practice (PSP), which keeps PI in the prompt and uses it only during sampling, not in the loss. PSP consistently outperforms plain GRPO across AppWorld and SWE‑bench Verified, improving task‑goal completion by up to 65% and resolved rate by up to 61%.

arXiv AI
Jul 7

TREK: Distill to Explore, Reinforce to Refine

arXiv:2607. 05339v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) is effective when the current policy already samples useful reasoning trajectories, but it stalls on hard prompts whose correct solution modes lie outside the student's on-policy support.

By Yuanda Xu, Zhengze Zhou, Kayhan Behdin, Jelena Markovic-Voronov, Hejian Sang, Xiaomin Li, Wenhui Zhu, Xinchen Du, Aida Rahmattalabi, Ran He, Sen Na, Zhipeng Wang, Alborz Geramifard