arXiv AI

WebRISE: Requirement-Induced State Evaluation for MLLM-Generated Web Artifacts

arXiv:2606. 03220v1 Announce Type: cross Abstract: Existing benchmarks for MLLM-generated web artifacts assess interaction through local evidence and miss the requirement-induced states and transitions that determine whether a page works.

arXiv AI
3d ago

WebUIProof: Benchmarking WebUI Code Generators with UI-Agent Execution Harness

WebUIProof is a new execution‑oriented benchmark for WebUI code generation that supplies structured specifications and dense, executable interaction tests across general WebUIs and 3D interactive simulations. It employs a UI‑agent harness that runs tests in a headless browser using a plan–act–observe loop to locate DOM elements, perform actions, observe changes, and verify assertions. Evaluations on eight commercial LLMs reveal frequent interaction‑based failures, especially on 3D interfaces, and demonstrate that training compact models with RL rewards from these tests improves functional completion and reduces build failures.

By Yun-Yun Tsai, Yuning Mao, Shiqi Wang, Junfeng Yang, Sinong Wang
arXiv Computer Vision
Sep 3

Rendering-in-the-Loop: An Execution-Driven Agent for Interactive Web Development

RILA is an execution‑driven agent that integrates browser rendering into the generation loop for interactive web development. It uses an Action Interaction Verification module to replay reference interactions on generated pages, collecting execution‑aware observations, and an Execution‑aware Rendering Score to jointly assess interaction correctness and visual fidelity during iterative optimization. A data synthesis pipeline further augments training data, enabling RILA to significantly improve interaction and visual quality across foundation models, even outperforming larger one‑shot generators.

By Yilong Guo, Hanqi Chen, Zixiao Ye, Guanzhong Wang, Chen Yu, Zeyu Chen
arXiv AI
Jun 17

LongWebBench: Evaluating Structural and Functional Webpage Generation in Long-Horizon Settings

arXiv:2606. 17727v1 Announce Type: new Abstract: Recent vision-language models (VLMs) have shown promising progress in generating webpages from visual inputs, yet existing evaluations mainly focus on short, single-screen, and largely static webpages.

By Yi Zhao, Zhen Yang, Mengpan Chen, Mingde Xu, Shanghui Gong, Xijun Liu, Jibing Gong, Jie Tang
arXiv AI
Sep 23

WebCraftBench: Evaluating Web Application Generation from a Software Testing Perspective

arXiv:2609.15387v3 Announce Type: replace-cross Abstract: Human evaluation provides a direct measure of the quality of LLM-generated web applications. However, fitting human judgments through automat...

By Chenxu Liu, Zilu Zou, Peizhong Gao, Jiawen Tao, Zhexin Zhang, Guang Chen, Haowei Lin, Ying Zhou, Tianyi Bai, Dolly Deng, Suncong Zheng, Maxm Pan
arXiv AI
Aug 20

ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents

ComponentBench is a new benchmark that evaluates computer‑use agents at the component level on modern web UIs. It contains 97 canonical UI components and 2,910 programmatically verified tasks, along with cleaned human reference trajectories for measuring task success and interaction efficiency. The benchmark also offers a scalable pipeline for auditing structural difficulty and synthesizing failure analyses across tasks and component families.

By Tianchen Guan, Xinlei Lin, Royce Cheng-Yue, Xiangjun Wang, Shuyan Zhou
arXiv AI
Aug 20

ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks

ChainWorld is a new benchmark that composes atomic OSWorld tasks into long‑horizon desktop workloads, creating 347 chains of length two to four that compare two renderings of the same task sequence. In single‑turn evaluation all tasks are presented together in one prompt, while in multi‑turn evaluation tasks are revealed one at a time. Across four current computer‑use agents, maximum chain completion is 31 %, with multi‑turn evaluation improving completion for three models but both protocols remaining challenging and exposing different failure profiles.

By Vincent Siu, Manasi Sharma, Dawn Song, Daniel Yue Zhang, Chenguang Wang, Ying Liu