arXiv AI

LiveEvalBench: Toward Open-World Evaluation for Web Generation

arXiv:2608. 03689v1 Announce Type: new Abstract: Large language models are increasingly capable of synthesizing executable frontend projects, yet existing benchmarks still treat web generation as a static evaluation problem.

arXiv AI
Sep 23

WebCraftBench: Evaluating Web Application Generation from a Software Testing Perspective

arXiv:2609.15387v3 Announce Type: replace-cross Abstract: Human evaluation provides a direct measure of the quality of LLM-generated web applications. However, fitting human judgments through automat...

By Chenxu Liu, Zilu Zou, Peizhong Gao, Jiawen Tao, Zhexin Zhang, Guang Chen, Haowei Lin, Ying Zhou, Tianyi Bai, Dolly Deng, Suncong Zheng, Maxm Pan
arXiv AI
6d ago

The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge

The paper introduces KNOWS, a benchmark for evaluating web agents that act as assistants by retrieving, synthesizing, and presenting information across complex, multi-step browser tasks. It outlines a task design rubric, evaluation protocol combining deterministic checks with LLM judgments, and reports that current agents achieve only modest success, with the best performing agent succeeding on less than 3% of tasks. The study highlights significant gaps in agents’ tool use, visual understanding, and long‑horizon reasoning.

By Alexander Gill, Md Farhan Ishmam, Xuyen Nguyen, Neha Bhat, Parker Henry DeYoung, Fateme Hashemi Chaleshtori, Nathan Stringham, Kenneth Marino, Ana Marasovi\'c
arXiv Computation and Language
Sep 23

WatchPoint: Executable User Feedback for Real-World Agentic Web Development

WatchPoint is a simulated‑user system that generates and runs diagnostic scripts against a live web application, producing structured observations to guide coding model retries. Unlike prior methods that rely on screenshots or non‑executable metrics, WatchPoint operates on Web‑Bench—a benchmark of 50 multi‑file web projects with 1,000 sequential tasks verified by deterministic end‑to‑end tests. It recovers 57.6% of diagnosed tasks, matching a human tester’s 54.5% recovery rate, and identifies when such simulated feedback is beneficial or should be withheld.

By Guanqun Yang, Wei Yang, Xueqing Liu
arXiv Computation and Language
Sep 25

An Empirical Study of Automating Agent Evaluation

The paper presents EvalAgent, an AI assistant that automates agent evaluation by encoding domain expertise into evaluation skills such as procedural instructions, reusable code, and dynamic API retrieval. EvalAgent constructs a trace-based pipeline that outputs metrics, executable code, and reports, and is evaluated using a new meta-evaluation framework and AgentEvalBench. Results show that EvalAgent improves the Eval@1 metric from 17.5% to 65% and receives 79.5% human expert preference, while ablation studies confirm the importance of evaluation skills.

By Kang Zhou, Sangmin Woo, Haibo Ding, Kiran Ramnath, Subramanian Chidambaram, Aosong Feng, Vinayak Arannil, Muhyun Kim, Ishan Singh, Darren Wang, Zhichao Xu, Megha Gandhi, Nirmal Prabhu, Soumya Smruti Mishra, Smeet Dhakecha, Vivek Singh, Gouri Pandeshwar, Lin Lee Cheong
arXiv AI
Jun 8

SW-$A^2$-Bench: Benchmarking Autonomous Software Agent Generation for Agentic Web

arXiv:2604. 04226v2 Announce Type: replace-cross Abstract: The Agentic Web is emerging as a paradigm in which autonomous software agents interact with online resources and with each other to accomplish user goals.

By Linyao Chen, Bo Huang, Qinlao Zhao, Shuai Shao, Zhi Han, Zicai Cui, Ziheng Zhang, Guangtao Zeng, Wenzheng Tang, Yikun Wang, Yuanjian Zhou, Zimian Peng, Yong Yu, Weiwen Liu, Hiroki Kobayashi, Weinan Zhang
arXiv AI
3d ago

E2E-SWE: Benchmarking LLMs on Building Working Codebases from Scratch

E2E-SWE is a benchmark that tests large language models’ ability to create complete, functional software repositories from scratch. It includes 186 tasks across 11 programming languages, each requiring an agent to build an installable project based solely on a natural‑language specification and an empty workspace, while passing a hidden test suite. The benchmark was crafted by software engineers and LLMs, then refined through iterative verification by autonomous agents to ensure clarity and solvability.

By Hantian Ding, Chloe Bi, Jiacheng Zhu, John Yang, Matt Deitke, Pengcheng Yin, Zijian Wang, Rui Hou