arXiv AI By Linh-An Phan, MingXue Wang, Guangyu Wu, Feng Pan, Zhaoyu Pang, Yanbin Zhang

Lightweight, Rubric-Guided Trajectory Evaluation for Production AI Agents

Read the original on arXiv AI →

LiteTrajEval is a lightweight architecture designed for budget‑bounded trajectory evaluation of large‑language‑model agents. It creates compact domain‑specific rule profiles offline, preprocesses trajectories online to flag heuristic failure signals, and then uses a single rubric‑guided LLM judge to generate structured diagnostic reports. On public Magnetic‑One‑style and τ‑bench‑style datasets, LiteTrajEval improves failure‑localization alignment with human annotations by 20–35 percentage points and reduces cost and evaluation time by 6× and 8× respectively, and has been deployed in an enterprise agentic platform.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 7

TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories

arXiv:2608. 06346v1 Announce Type: new Abstract: LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging.

By Yunjia Qi, Zehua Yin, Xintong Shi, Hao Peng, Songyuanyi Lu, Yixian Liu, Richeng Xuan, Yuhong Liu, Zhichao Hu, Xiaozhi Wang, Lei Hou, Bin Xu, Juanzi Li
arXiv AI
Jul 16

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

arXiv:2607. 13705v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical.

By Zichen Ding, Jiaye Ge, Shufan Jiang, Kai Chen, Mo Li, Qingqiu Li, Zehao Li, Zonglin Li, Tiaohao Liang, Shudong Liu, Zerun Ma, Zixing Shang, Wenhui Tian, Zun Wang, Liwei Wu, Zhenyu Wu, Jun Xu, Bowen Yang, Dingbo Yuan, Qi Zhang, Songyang Zhang, Peiheng Zhou, Dongsheng Zhu