arXiv:2607. 19375v1 Announce Type: cross Abstract: Language models perform economically valuable work, yet they are not currently assessed for how well they perform every economically valuable task.
By Alexander Wan, Stephane Hatgis-Kessell, Tom\'as Aguirre, Percy Liang, Rishi Bommasani
The paper introduces the Agentic Adoption Index (AAI), a new measure of delegated exposure that captures whether workers actually commit tasks to AI within structured workflows. Using semantic embeddings of 888,000 agent skill specifications from GitHub and 18,000 O*NET task statements, the authors find that occupations with high delegation differ from those most vulnerable to pre-AI automation, that AAI correlates more with technical capability than with current LLM use, and that for lower‑educated occupations AAI rises with wages while it falls for higher‑educated, high‑earning workers. These patterns also appear in an independent corpus from the Manus Skills Marketplace.
By Hyeongjae Lee, Jihyang Cheon, Lanu Kim
The paper introduces the Agentic Adoption Index (AAI), a new metric that captures whether workers actually delegate tasks to AI within their workflows, rather than merely measuring potential AI applicability. Using 53,000 agent skill specifications and 18,000 O*NET task statements, the authors find that occupations with high delegation differ from those previously deemed most at risk, that AAI aligns more closely with AI’s capabilities than current usage, and that adoption peaks at mid‑wage, bachelor’s‑level occupations while declining at both ends of the wage and education spectrum. The study highlights that technical availability explains much of the variation, but other factors—such as resistance to specification or professional discretion—also influence who adopts AI.
whyItMatters":"The findings suggest that actual AI adoption patterns differ from prior risk assessments, indicating that factors beyond technical feasibility shape who delegates to AI, which has implications for workforce planning and policy."
By Hyeongjae Lee, Jihyang Cheon, Lanu Kim
arXiv:2606. 07489v1 Announce Type: new Abstract: Frontier AI systems are bridging the gap between intelligence and utility by shifting from conversational assistants to autonomous agents that execute tasks end to end.
By Jeremy Yang, Kate Zyskowski, Noah Yonack, Jerry Ma
The paper introduces the Token Economy Score (TES), a metric that quantifies the accuracy gain of reasoning-capable large language models relative to non-reasoning baselines, normalized by token generation cost. An empirical study across 151 runs on seven diverse benchmarks shows that task structure—such as sequential inference chains—predicts higher TES, while knowledge-recall tasks yield lower TES despite difficulty. The analysis also reveals diminishing returns at higher reasoning effort and highlights how deployment context, via Reasoning Cost Share and Deployment Cost Multiplier, can alter the economic viability of reasoning workloads.
By Sachin Gopal Wani, Ajay Dholakia, David Ellison
arXiv:2608. 00355v1 Announce Type: cross Abstract: Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score.
By Hanwen Xing, Pengyun Wang, BingXu Meng, Kumail Alhamoud, Xiang Li, Jicheng Wang, Xin Yu, Xinyang Han, Xiaomin Li, Philip Torr, Yuexing Hao