arXiv AI By Qitai Tan, Zefang Zong, Mo Li, Yipeng Shi, Yang Li, Peng Chen

ATOD: Annealed Turn-Aware On-Policy Distillation for Multi-Turn Agentic Tasks

Read the original on arXiv AI →

arXiv:2606. 27814v4 Announce Type: replace Abstract: Training small language-model agents for long-horizon interactive tasks requires both fast imitation and reward-driven improvement.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.