arXiv AI By Andrew Zhang, Chengzhan Li

Agent Step Value: State-Transition Measurement with State-Grounded LLM Evaluators

Read the original on arXiv AI →

arXiv:2607. 04419v1 Announce Type: new Abstract: Most agent evaluations collapse a multi-step trace into a final answer, a success flag, or a trajectory-level score.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 31

ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents

arXiv:2607. 28037v1 Announce Type: new Abstract: As LLM-based agents are deployed in complex, multi-step workflows, a critical evaluation gap has emerged: most existing benchmarks judge only final outcomes, unable to distinguish reliable reasoning from lucky success or attribute failures to specific process deficiencies, hindering attribution in long-horizon tasks.

By Xingjian Wu, Xuhang Zhu, Xingchen Liu, Junlin Liu, Jianing Wang, Linsen Guo, Xiaoyu Li, Xuezhi Cao, Xunliang Cai