arXiv Machine Learning By Chengheng Li-Chen, Zhiqian Zhou, Hao Chen, Nicolas Chauvin

WinDOM: Self-Family Distillation for Small-Model GUI Grounding

Read the original on arXiv Machine Learning →

arXiv:2606. 25964v1 Announce Type: cross Abstract: Small ($\sim$2B) GUI-grounding agents are attractive for on-device deployment, accessibility tooling, and low-cost iteration, but at this scale they face two open recipe questions: how to obtain bounding-box training data without expensive human annotation, and how to combine supervised fine-tuning with reinforcement learning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 25

From Self-Distillation to Self-Practice: Privileged Information for Multi-Turn Agents

The paper introduces Privileged Self-Practice (PSP), a method that retains privileged information (PI) in the prompt rather than the loss during on‑policy self‑distillation for multi‑turn agents. PSP injects short per‑task instructions from an analyzer model when rollouts fail, sampling again with the instruction in context and training with the unchanged GRPO objective. Experiments on AppWorld and SWE‑bench Verified show PSP consistently outperforms plain GRPO, boosting task‑goal completion by up to 65% and resolved rate by up to 61% across three student models.

By Xingyu Su, Abhishek Kumar, Qing Ping, Youzhi Luo, Jonathan Buck, Zach Zhang, Subramanian Chidambaram, Vinayak Arannil