arXiv AI
2d ago

ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents

ComputerSD is an online self‑distillation method for computer‑use agents that leverages real‑time feedback from executed GUI transitions. It uses a fine‑tuned GUI analyzer to generate guidance and a step‑level value score after each action, combining token‑level OPSD with trajectory‑level GRPO in an asynchronous training framework. On the OSWorld‑Verified benchmark, ComputerSD improves performance over outcome‑only GRPO by 1.9 and 4.1 percentage points on Qwen3‑VL‑8B‑Thinking and EvoCUA‑8B backbones, and shows strong generalizability in out‑of‑distribution tests.

By Yong Du, Tongbo Chen, Zhengxi Lu, Yizhou Liu, Bofan Chen, Tao Jiang, Wenhao Xu, Yongliang Shen
Hugging Face Trending Papers
Aug 13

Latent On-Policy Self-Distillation

Enabling agents to learn from experience and internalize it into their policy has become a central problem in self-evolving AI. On-policy self-distillation (OPSD) offers an effective pathway by using a privileged self-teacher to provide dense supervision on the student's own trajectories; however, existing methods still rely heavily on designer-specified privileged artifacts (e.