WorkWorlds: An Infrastructure for Evaluating AI Agents on Workplace Tasks
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper proposes a new way to describe benchmarks for AI systems that perform knowledge work, outlining four explicit fields: represented activity, tested setting, required work product, and evaluated result. It builds an inventory of 18 work activities from O*NET to enable activity-level reporting across occupations, and evaluates these activities for semantic coherence, algorithm sensitivity, ontology legibility, and human interpretability. The authors apply their framework to three existing benchmarks—GDPval, OfficeQA Pro, and APEX-SWE—to show how different aspects of work can be captured within the same representation.
arXiv:2606. 05405v1 Announce Type: cross Abstract: Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional domains.
The paper introduces KNOWS, a benchmark for evaluating web agents that act as assistants by retrieving, synthesizing, and presenting information across complex, multi-step browser tasks. It outlines a task design rubric, evaluation protocol combining deterministic checks with LLM judgments, and reports that current agents achieve only modest success, with the best performing agent succeeding on less than 3% of tasks. The study highlights significant gaps in agents’ tool use, visual understanding, and long‑horizon reasoning.
arXiv:2609.39325v1 Announce Type: new Abstract: The ability of Large Language Model (LLM) agents to complete daily and professional work is receiving increasing attention. Training such agents requir...
arXiv:2607. 25398v1 Announce Type: new Abstract: Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let it govern every action that follows.
DAYJOB is a new benchmark comprising 130 long‑horizon professional tasks created by experts in healthcare (50 tasks) and finance (80 tasks). Each task is packaged as a containerized Harbor environment and evaluated against a detailed binary rubric, requiring agents to meet every criterion to pass. In tests across 30 model configurations, the best model (Claude Opus 5.5) achieves about 24% success in both domains, while the median model scores below 3%.