arXiv AI By Kaibing Yang, Guangfeng Cai, Shengtian Yang, Shuo He, Yu Li, Mengyi Liu, Pengwei Chen, Jun Xu, Lei Feng

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks

Read the original on arXiv AI →

arXiv:2607. 22724v1 Announce Type: cross Abstract: Group-based policy optimization has been increasingly used to train large language model (LLM) agents from sparse outcome rewards by comparing trajectories or steps within a group.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.