Economic Evaluations of Language Models
arXiv:2607. 19375v1 Announce Type: cross Abstract: Language models perform economically valuable work, yet they are not currently assessed for how well they perform every economically valuable task.
arXiv:2608. 05172v1 Announce Type: cross Abstract: The task-based framework in economics models occupations as bundles of tasks.
arXiv:2607. 19375v1 Announce Type: cross Abstract: Language models perform economically valuable work, yet they are not currently assessed for how well they perform every economically valuable task.
The paper introduces the Agentic Adoption Index (AAI), a new measure of delegated exposure that captures whether workers actually commit tasks to AI within structured workflows. Using semantic embeddings of 888,000 agent skill specifications from GitHub and 18,000 O*NET task statements, the authors find that occupations with high delegation differ from those most vulnerable to pre-AI automation, that AAI correlates more with technical capability than with current LLM use, and that for lower‑educated occupations AAI rises with wages while it falls for higher‑educated, high‑earning workers. These patterns also appear in an independent corpus from the Manus Skills Marketplace.
The paper introduces the Agentic Adoption Index (AAI), a new metric that captures whether workers actually delegate tasks to AI within their workflows, rather than merely measuring potential AI applicability. Using 53,000 agent skill specifications and 18,000 O*NET task statements, the authors find that occupations with high delegation differ from those previously deemed most at risk, that AAI aligns more closely with AI’s capabilities than current usage, and that adoption peaks at mid‑wage, bachelor’s‑level occupations while declining at both ends of the wage and education spectrum. The study highlights that technical availability explains much of the variation, but other factors—such as resistance to specification or professional discretion—also influence who adopts AI. whyItMatters":"The findings suggest that actual AI adoption patterns differ from prior risk assessments, indicating that factors beyond technical feasibility shape who delegates to AI, which has implications for workforce planning and policy."
arXiv:2606. 07489v1 Announce Type: new Abstract: Frontier AI systems are bridging the gap between intelligence and utility by shifting from conversational assistants to autonomous agents that execute tasks end to end.
The paper introduces the Token Economy Score (TES), a metric that quantifies the accuracy gain of reasoning-capable large language models relative to non-reasoning baselines, normalized by token generation cost. An empirical study across 151 runs on seven diverse benchmarks shows that task structure—such as sequential inference chains—predicts higher TES, while knowledge-recall tasks yield lower TES despite difficulty. The analysis also reveals diminishing returns at higher reasoning effort and highlights how deployment context, via Reasoning Cost Share and Deployment Cost Multiplier, can alter the economic viability of reasoning workloads.
arXiv:2608. 00355v1 Announce Type: cross Abstract: Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score.
A set of exposure scores calculated in 2023 has become a central empirical input to the future of work debate. Produced by Eloundou et al.
arXiv:2602. 07267v2 Announce Type: replace Abstract: Evaluating the real-world capabilities of AI systems requires grounding benchmark performance in human-interpretable measures of task difficulty.
arXiv:2606. 04402v1 Announce Type: new Abstract: Modern reasoning models can allocate different amounts of test-time computation, such as thinking tokens, model calls, or compute budget, to different tasks.
arXiv:2607. 06283v1 Announce Type: new Abstract: Skill usage can significantly enhance the ability of modern agent systems to complete complex tasks.
The paper introduces a low‑rank framework for ranking large language models (LLMs) on task‑specific benchmarks using sparse pairwise comparisons. By modeling the task‑by‑model ability matrix as low rank, the method shares information across related tasks while preserving task‑specific differences, and it provides uncertainty‑aware ranking through debiased estimators and simultaneous confidence sets. Experiments on synthetic data and the Chatbot Arena benchmark demonstrate improved sample efficiency and tighter, better‑calibrated ranking certificates, especially in the sparse comparison regime typical of real LLM evaluations.
SCHEDBench is a natural‑language benchmark that evaluates whether large language models (LLMs) produce schedules that remain constraint‑feasible when the same scheduling problem is expressed in different natural‑language surface forms. The benchmark covers 1,132 instances from job‑shop scheduling, resource‑constrained project scheduling, nurse rostering, and curriculum timetabling, and uses domain‑specific templates and surface‑form variations to generate varied problem statements. Experiments with thirteen frontier and open‑weight LLMs show that models are not reliably invariant to semantically equivalent renderings, with surface‑form variation reducing feasibility and increasing hard‑constraint violations, especially when constraints are reordered.