arXiv AI

Multi-Agent LLM Orchestration Achieves Deterministic, High-Quality Decision Support for Incident Response

arXiv AI
Aug 28

DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows

DuMateBench is a new benchmark for autonomous agents that uses real user sessions from a large production platform, preserving interaction history, configurations, and workspace state. It contains 200 tasks across 8 scenarios and 17 capability categories, many requiring coordination of multiple capabilities. The benchmark tests agents in Docker containers with real-world complexities—Insufficient, Unstable, and Noisy—and evaluates performance with a hybrid deterministic and LLM-as-Judge protocol, revealing significant gaps in task completion across various agent frameworks and LLMs.

By Zechun Niu, Yukun Zhao, Jiaxin Zhang, Xu Shen, Jinhua Si, Han Tian, Can Xu, Yunfan Song, Jiaxin Mao, Yansong Gao, Yuchen Li, Jianmin Wu, Lingyong Yan, Shuaiqiang Wang, Dawei Yin
arXiv AI
Jul 14

AgentAbstain: Do LLM Agents Know When Not to Act?

arXiv:2607. 10059v1 Announce Type: new Abstract: Agent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain.

By Xun Liu, Yi Evie Zhang, Vira Kasprova, Parisa Rabbani, Pardis Sadat Zahraei, Tianyu Zhang, Ali Ebrahimpour-Boroojeny, Varun Chandrasekaran