arXiv AI By Qitai Tan, Zefang Zong, Yang Li, Peng Chen

ATOD: Annealed Turn-aware On-policy Distillation for Multi-turn Autonomous Agents

Read the original on arXiv AI →

arXiv:2606. 27814v1 Announce Type: new Abstract: Training small language-model agents for long-horizon interactive tasks requires both fast imitation and reward-driven improvement.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.