ClawBench: Can AI Agents Complete Everyday Online Tasks?
arXiv:2604. 08523v2 Announce Type: replace-cross Abstract: AI agents may be able to assist with emails and documents, but can they reliably complete everyday online workflows on real websites?
arXiv:2606. 16748v1 Announce Type: new Abstract: Current benchmarks for computer-use agents evaluate models in impersonal environments.
arXiv:2604. 08523v2 Announce Type: replace-cross Abstract: AI agents may be able to assist with emails and documents, but can they reliably complete everyday online workflows on real websites?
arXiv:2604. 13072v2 Announce Type: replace-cross Abstract: OpenClaw-style personal assistants extend LLM agents from isolated tool use to open-ended, stateful, and personalized software environments.
arXiv:2606. 28480v1 Announce Type: cross Abstract: As large language models and harness frameworks continue to advance, agents operating in terminals are increasingly capable of performing a broader range of general computer-use tasks beyond coding.
RealClawBench is a live benchmark framework derived from real OpenClaw developer‑agent sessions, designed to capture the distribution, diversity, and real‑world difficulty of deployed agent use. It reconstructs execution environments and uses deterministic verifiable scorers to convert real sessions into reproducible, automatically scored tasks, yielding 281 executable tasks with minimal distribution shift. Evaluation of 14 contemporary models shows the best system solves only 65.8% of tasks, highlighting significant room for improvement on realistic workloads.
arXiv:2506. 01952v2 Announce Type: replace-cross Abstract: Powered by large language models (LLMs), web browsing agents operate graphical user interfaces in a human-like manner, offering a transparent and general framework for automating web-based tasks.
arXiv:2606. 10394v1 Announce Type: new Abstract: Large language models are increasingly used to power personal agents for everyday applications, but evaluating these agents remains a challenge.
arXiv:2606. 29537v2 Announce Type: replace Abstract: Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world computer use, limiting their ability to reveal the limitations of frontier agents.
arXiv:2606. 29537v1 Announce Type: new Abstract: Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world computer use, limiting their ability to reveal the limitations of frontier agents.
Wuying-Browser-Agent is a unified framework designed to improve long-horizon browser agents by aligning execution, supervision, optimization, and evaluation. It introduces a structured browser harness, reflection and UI-specialized Curriculum SFT (RUIC‑SFT) for recovery and complex UI interactions, and Divergence‑Aware Online GRPO (DAO‑GRPO) for better credit assignment. The framework is evaluated on BrowserBench—a bilingual real‑web benchmark of 350 tasks—and achieves state‑of‑the‑art results on multiple browser‑use benchmarks, while also transferring well to other agentic tasks.
Large language models are increasingly used to power personal agents for everyday applications, but evaluating these agents remains a challenge. Existing benchmarks still rely on sandboxed artifacts, static task design, and coarse scoring, which hinder scalability and limit progress toward reliable personal-agent evaluation.
arXiv:2609.35814v1 Announce Type: cross Abstract: As browser-use agents improve, benchmarks keep pace by collecting new tasks, websites, and applications, often making tasks longer or more novel. Thi...
Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments.