IWC-Bench: Evaluating Web Application Generation from a Software Testing Perspective
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2609.15387v3 Announce Type: replace-cross Abstract: Human evaluation provides a direct measure of the quality of LLM-generated web applications. However, fitting human judgments through automat...
arXiv:2601. 02430v3 Announce Type: replace-cross Abstract: Web applications (web apps) have become a key arena for large language models (LLMs) to demonstrate their code generation capabilities and commercial potential.
arXiv:2608. 03689v1 Announce Type: new Abstract: Large language models are increasingly capable of synthesizing executable frontend projects, yet existing benchmarks still treat web generation as a static evaluation problem.
WatchPoint is a simulated‑user system that generates and runs diagnostic scripts against a live web application, producing structured observations to guide coding model retries. Unlike prior methods that rely on screenshots or non‑executable metrics, WatchPoint operates on Web‑Bench—a benchmark of 50 multi‑file web projects with 1,000 sequential tasks verified by deterministic end‑to‑end tests. It recovers 57.6% of diagnosed tasks, matching a human tester’s 54.5% recovery rate, and identifies when such simulated feedback is beneficial or should be withheld.
arXiv:2605. 30000v2 Announce Type: replace Abstract: Front-end web code has become a core product surface for every frontier LLM release, yet evaluating these interactive applications at development speed remains costly because human-judged leaderboards like Arena do not scale.
arXiv:2605. 26144v2 Announce Type: replace-cross Abstract: We present VISTA (VIsual Spec-To-App Benchmark), a benchmark for evaluating the end-to-end web-app generation capabilities of LLM-based agents.