arXiv AI By Zihang Tian, Jingsen Zhang, Rui Li, Xiaohe Bo, Yuanzi Li, Xu Chen

ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents

Read the original on arXiv AI →

arXiv:2606. 21262v2 Announce Type: replace Abstract: Reinforcement learning for multi-step LLM agents often relies on scalar rewards that indicate success but cannot explain why a trajectory is good or bad.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.