arXiv:2606. 16748v1 Announce Type: new Abstract: Current benchmarks for computer-use agents evaluate models in impersonal environments.
By Lawrence Keunho Jang, Andrew Keunwoo Jang, Jing Yu Koh, Ruslan Salakhutdinov
arXiv:2608.22510v1 Announce Type: new
Abstract: Agent benchmarks often evaluate only final answers even when agents run on stateful runtimes. We argue this under-specifies what is being evaluated: th...
By YuanHang Xiao
arXiv:2604. 13072v2 Announce Type: replace-cross Abstract: OpenClaw-style personal assistants extend LLM agents from isolated tool use to open-ended, stateful, and personalized software environments.
By Xiang Long, Li Du, Yilong Xu, RongJian Xu, Qiyanhui Lu, Ying Gao, Qinhua Xie, Fangcheng Liu, Ning Ding, Haoqing Wang, Ziheng Li, Changjiang Zhou, Jianyuan Guo, Yehui Tang
arXiv:2609.06059v1 Announce Type: new
Abstract: As large language models evolve from question-answering systems into general-purpose agents, evaluation must move beyond static answer correctness to a...
By Yu Liu, Zhilin Liu, Zhiwei Yang, Shaojie Zhang, Zheyuan Deng, Tingwei Huang, Zhenbo Luo, Lei Jiang, Yanbing Liu, Pei Fu
StartupBench is a new benchmark that evaluates general-purpose agents on end-to-end workflows derived from AI startup products that have proven market adoption. It translates real-world product workflows into deliverable-oriented tasks and assesses them with detailed rubrics. The study finds that even the best models complete only about 30% of these tasks, highlighting challenges such as complex instruction following and domain expertise.
By Liya Zhu, Xin Ma, Tao Liu, Haodong Wang, Ge Zhang, Jingzhe Ding, Qingshui Gu, Yongjie Zhong, Jinxiang Meng, Yuan Gao, Yunqiu Zhou, Hao Zhu, Jifeng He, Yongzhi Liao, Xinyi Zhang, Chaoxin Li, Yi Zhu, Xi Lin, Duju Zeng, Xiang Gao, Wen Zhang, Yunyang Wang, Duo Wang, Huan Zhou, Zuo Wang, Jin Chen, Kaiyuan Zhang, Chuqian Yu, Tianhao Yu, Longxiang Liu, Jianbo Xue, Huimin Che, Jiahao Wang, Yujia Qin, Jiaheng Liu, Shen Yan, Xiaolong Chang, Wenhao Huang
StartupBench is a benchmark that evaluates general‑purpose agents on end‑to‑end workflows derived from real AI startup products that have proven market adoption. It translates these product workflows into deliverable‑oriented tasks and assesses them with detailed rubrics that capture complex requirements. Even the best current models complete only about 30% of the tasks, highlighting failures in instruction following and domain expertise.