MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents
arXiv:2606. 16748v1 Announce Type: new Abstract: Current benchmarks for computer-use agents evaluate models in impersonal environments.
arXiv:2604. 08523v2 Announce Type: replace-cross Abstract: AI agents may be able to assist with emails and documents, but can they reliably complete everyday online workflows on real websites?
arXiv:2606. 16748v1 Announce Type: new Abstract: Current benchmarks for computer-use agents evaluate models in impersonal environments.
arXiv:2608.22510v1 Announce Type: new Abstract: Agent benchmarks often evaluate only final answers even when agents run on stateful runtimes. We argue this under-specifies what is being evaluated: th...
arXiv:2604. 13072v2 Announce Type: replace-cross Abstract: OpenClaw-style personal assistants extend LLM agents from isolated tool use to open-ended, stateful, and personalized software environments.
arXiv:2609.06059v1 Announce Type: new Abstract: As large language models evolve from question-answering systems into general-purpose agents, evaluation must move beyond static answer correctness to a...
StartupBench is a new benchmark that evaluates general-purpose agents on end-to-end workflows derived from AI startup products that have proven market adoption. It translates real-world product workflows into deliverable-oriented tasks and assesses them with detailed rubrics. The study finds that even the best models complete only about 30% of these tasks, highlighting challenges such as complex instruction following and domain expertise.
StartupBench is a benchmark that evaluates general‑purpose agents on end‑to‑end workflows derived from real AI startup products that have proven market adoption. It translates these product workflows into deliverable‑oriented tasks and assesses them with detailed rubrics that capture complex requirements. Even the best current models complete only about 30% of the tasks, highlighting failures in instruction following and domain expertise.
arXiv:2606. 07682v1 Announce Type: cross Abstract: AI agents are increasingly expected to complete long-horizon workflows that require sustained progress over hours, millions of tokens, and complex environments.
arXiv:2606. 28480v1 Announce Type: cross Abstract: As large language models and harness frameworks continue to advance, agents operating in terminals are increasingly capable of performing a broader range of general computer-use tasks beyond coding.
arXiv:2606. 05342v1 Announce Type: new Abstract: AI agents are increasingly asked to carry out work that spans minutes, hours, or longer.
arXiv:2606. 12871v1 Announce Type: new Abstract: Search Agents (SAs) typically leverage large language models (LLMs) to support complex information-seeking tasks by autonomously exploring web sources and synthesizing information into comprehensive responses.
arXiv:2606. 09426v1 Announce Type: new Abstract: Computer-use agents (CUAs) increasingly operate in runtimes that combine visual desktop control, command-line execution, code editing, browsers, and external tools.
Wuying-Browser-Agent is a unified framework designed to improve long-horizon browser agents by aligning execution, supervision, optimization, and evaluation. It introduces a structured browser harness, reflection and UI-specialized Curriculum SFT (RUIC‑SFT) for recovery and complex UI interactions, and Divergence‑Aware Online GRPO (DAO‑GRPO) for better credit assignment. The framework is evaluated on BrowserBench—a bilingual real‑web benchmark of 350 tasks—and achieves state‑of‑the‑art results on multiple browser‑use benchmarks, while also transferring well to other agentic tasks.