arXiv AI By Hyun Bin Park (Sogang University), Kyungho Song (University of Michigan, Ann Arbor), Sangmin Lee (Sogang University), Du-Seong Chang (Sogang University)

Persistent Teacher Anchoring for Tool-Using Agents

Read the original on arXiv AI →

Persistent Teacher Anchoring (PTA) is a method that extends on‑policy knowledge distillation by ensuring that a teacher verifies entire turns before any tool calls are executed. PTA builds on chunk‑level verification with an added turn‑level commitment, treating verified chunks as atomic units and introducing persistent lookahead to keep rollout capacity full. Experiments on Search‑R1 and DeepEyes show that PTA improves macro best@4 by 2.5–2.8 points over standard OPKD and boosts throughput by 24%.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 19

SOD: Step-wise On-policy Distillation for Small Language Model Agents

SOD: Step-wise On-policy Distillation for Small Language Model Agents proposes a new framework that adaptively reweights distillation strength at each reasoning step based on step-level divergence. This approach mitigates cascading errors in tool-integrated reasoning by attenuating misleading teacher signals in high-divergence regions while preserving dense guidance where student and teacher align. Experiments on math, science, and code benchmarks show up to 20.86% improvement over the second-best baseline, with a 0.6B student scoring 26.13% on AIME 2025.

By Qiyong Zhong, Mao Zheng, Mingyang Song, Xin Lin, Jie Sun, Houcheng Jiang, Xiang Wang, Junfeng Fang
arXiv AI
Jun 9

Trajectory-Refined Distillation

arXiv:2606. 08432v1 Announce Type: new Abstract: On-policy distillation (OPD) has become a central post-training tool for large language models (LLMs), providing dense per-token teacher supervision along the student's own rollouts.

By Li Jiang, Haoran Xu, Yichuan Ding, Amy Zhang