arXiv Machine Learning

Agents' Last Exam

arXiv:2606. 05405v1 Announce Type: cross Abstract: Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional domains.

arXiv AI
Aug 24

Who Delegates to AI? Evidence from 53,000 Agent Configurations

The paper introduces the Agentic Adoption Index (AAI), a new metric that captures whether workers actually delegate tasks to AI within their workflows, rather than merely measuring potential AI applicability. Using 53,000 agent skill specifications and 18,000 O*NET task statements, the authors find that occupations with high delegation differ from those previously deemed most at risk, that AAI aligns more closely with AI’s capabilities than current usage, and that adoption peaks at mid‑wage, bachelor’s‑level occupations while declining at both ends of the wage and education spectrum. The study highlights that technical availability explains much of the variation, but other factors—such as resistance to specification or professional discretion—also influence who adopts AI. whyItMatters":"The findings suggest that actual AI adoption patterns differ from prior risk assessments, indicating that factors beyond technical feasibility shape who delegates to AI, which has implications for workforce planning and policy."

By Hyeongjae Lee, Jihyang Cheon, Lanu Kim
arXiv AI
Aug 25

Designing Benchmarks for Knowledge Work

The paper proposes a new way to describe benchmarks for AI systems that perform knowledge work, outlining four explicit fields: represented activity, tested setting, required work product, and evaluated result. It builds an inventory of 18 work activities from O*NET to enable activity-level reporting across occupations, and evaluates these activities for semantic coherence, algorithm sensitivity, ontology legibility, and human interpretability. The authors apply their framework to three existing benchmarks—GDPval, OfficeQA Pro, and APEX-SWE—to show how different aspects of work can be captured within the same representation.

By Yining Hua, Hongbin Na, Cyrus Ayubcha, Levi Lian
arXiv AI
Jun 2

AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science

arXiv:2603. 19005v2 Announce Type: replace-cross Abstract: Data science plays a critical role in transforming complex data into actionable insights across numerous domains.

By An Luo, Jin Du, Xun Xian, Robert Specht, Fangqiao Tian, Ganghua Wang, Xuan Bi, Charles Fleming, Ashish Kundu, Jayanth Srinivasa, Mingyi Hong, Rui Zhang, Tianxi Li, Galin Jones, Jie Ding
arXiv AI
Jul 15

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks

arXiv:2606. 29537v2 Announce Type: replace Abstract: Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world computer use, limiting their ability to reveal the limitations of frontier agents.

By Mengqi Yuan, Zilong Zhou, Xinzhuang Xiong, Weiming Wu, Jiayang Sun, Jiamin Song, Kaiqian Cui, Bowen Wang, Haoyuan Wu, Yitong Li, Dunjie Lu, Haikong Lu, Qi Zhen, Xinyuan Wang, Jiaqi Deng, Yuhao Yang, Cheng Chen, Boyuan Zheng, Alex Su, Xiao Yu, Hao Zou, Saaket Agashe, Xing Han Lu, Manpreet Kaur, Zhengyang Qi, Vincent Sunn Chen, Frederic Sala, Dayiheng Liu, Junyang Lin, Zhou Yu, Yu Su, Siva Reddy, Xin Eric Wang, Peng Qi, Tianbao Xie, Tao Yu
arXiv AI
Jun 30

OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks

arXiv:2606. 29537v1 Announce Type: new Abstract: Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world computer use, limiting their ability to reveal the limitations of frontier agents.

By Mengqi Yuan, Zilong Zhou, Xinzhuang Xiong, Weiming Wu, Jiayang Sun, Jiamin Song, Kaiqian Cui, Bowen Wang, Haoyuan Wu, Yitong Li, Dunjie Lu, Haikong Lu, Qi Zhen, Xinyuan Wang, Jiaqi Deng, Yuhao Yang, Cheng Chen, Boyuan Zheng, Alex Su, Xiao Yu, Hao Zou, Saaket Agashe, Xing Han Lu, Manpreet Kaur, Zhengyang Qi, Vincent Sunn Chen, Frederic Sala, Dayiheng Liu, Junyang Lin, Zhou Yu, Yu Su, Siva Reddy, Xin Eric Wang, Peng Qi, Tianbao Xie, Tao Yu
arXiv AI
Sep 10

Who Delegates to AI? Evidence from Agent Configurations in Github

The paper introduces the Agentic Adoption Index (AAI), a new measure of delegated exposure that captures whether workers actually commit tasks to AI within structured workflows. Using semantic embeddings of 888,000 agent skill specifications from GitHub and 18,000 O*NET task statements, the authors find that occupations with high delegation differ from those most vulnerable to pre-AI automation, that AAI correlates more with technical capability than with current LLM use, and that for lower‑educated occupations AAI rises with wages while it falls for higher‑educated, high‑earning workers. These patterns also appear in an independent corpus from the Manus Skills Marketplace.

By Hyeongjae Lee, Jihyang Cheon, Lanu Kim
arXiv AI
Jun 17

A Framework for Evaluating Agentic Skills at Scale

arXiv:2606. 17819v1 Announce Type: cross Abstract: Agent skills -- structured, reusable knowledge artifacts that augment LLM agent capabilities -- have been rapidly adopted in industry, yet their cross-domain impact and use across commercial and open-source models remain under-studied, and no reusable methodology exists for evaluating an individual skill.

By Maksim Shaposhnikov, Nicolas Fortuin, Simon Stipcich, Maria I. Gorinova, Amy Heineike, Rob Willoughby