WebSP-Eval: Evaluating Web Agents on Website Security and Privacy Tasks
arXiv:2604. 06367v2 Announce Type: replace-cross Abstract: Web agents automate browser tasks, ranging from simple form completion to complex workflows like ordering groceries.
arXiv:2602. 17245v2 Announce Type: replace Abstract: This position paper argues that building a reliable agentic Web requires shifting from low-level interaction primitives to typed actions supported by a semantic layer.
arXiv:2604. 06367v2 Announce Type: replace-cross Abstract: Web agents automate browser tasks, ranging from simple form completion to complex workflows like ordering groceries.
arXiv:2608. 08392v1 Announce Type: new Abstract: Large language models are increasingly deployed as autonomous agents that interact with the web through browsers.
The paper introduces KNOWS, a benchmark for evaluating web agents that act as assistants by retrieving, synthesizing, and presenting information across complex, multi-step browser tasks. It outlines a task design rubric, evaluation protocol combining deterministic checks with LLM judgments, and reports that current agents achieve only modest success, with the best performing agent succeeding on less than 3% of tasks. The study highlights significant gaps in agents’ tool use, visual understanding, and long‑horizon reasoning.
arXiv:2606. 14027v1 Announce Type: cross Abstract: Agentic browsers integrate autonomous AI agents into web browsers, enabling users to accomplish web tasks through natural-language instructions.
arXiv:2608.21898v1 Announce Type: new Abstract: Web agents promise to automate complex digital workflows, but their training remains limited by synthetic environments that look plausible while hiding...
arXiv:2510. 19838v2 Announce Type: replace Abstract: Autonomous web agents powered by large language models (LLMs) show strong potential for performing goal-oriented tasks such as information retrieval, report generation, and online transactions.
arXiv:2506. 01952v2 Announce Type: replace-cross Abstract: Powered by large language models (LLMs), web browsing agents operate graphical user interfaces in a human-like manner, offering a transparent and general framework for automating web-based tasks.
arXiv:2609.35814v1 Announce Type: cross Abstract: As browser-use agents improve, benchmarks keep pace by collecting new tasks, websites, and applications, often making tasks longer or more novel. Thi...
arXiv:2606. 15673v1 Announce Type: new Abstract: Web agents act through long interaction sequences, yet existing benchmarks evaluate only terminal success, discarding all process information and offering little guidance on improvement.
arXiv:2608.28597v1 Announce Type: new Abstract: Online surveys are a foundational data collection instrument in a variety of fields, with attention checks serving as critical guardians of response qu...
arXiv:2608. 03689v1 Announce Type: new Abstract: Large language models are increasingly capable of synthesizing executable frontend projects, yet existing benchmarks still treat web generation as a static evaluation problem.
WatchPoint is a simulated‑user system that generates and runs diagnostic scripts against a live web application, producing structured observations to guide coding model retries. Unlike prior methods that rely on screenshots or non‑executable metrics, WatchPoint operates on Web‑Bench—a benchmark of 50 multi‑file web projects with 1,000 sequential tasks verified by deterministic end‑to‑end tests. It recovers 57.6% of diagnosed tasks, matching a human tester’s 54.5% recovery rate, and identifies when such simulated feedback is beneficial or should be withheld.