arXiv AI

Estimating time spent on work tasks

arXiv:2608. 05172v1 Announce Type: cross Abstract: The task-based framework in economics models occupations as bundles of tasks.

arXiv AI
Sep 10

Who Delegates to AI? Evidence from Agent Configurations in Github

The paper introduces the Agentic Adoption Index (AAI), a new measure of delegated exposure that captures whether workers actually commit tasks to AI within structured workflows. Using semantic embeddings of 888,000 agent skill specifications from GitHub and 18,000 O*NET task statements, the authors find that occupations with high delegation differ from those most vulnerable to pre-AI automation, that AAI correlates more with technical capability than with current LLM use, and that for lower‑educated occupations AAI rises with wages while it falls for higher‑educated, high‑earning workers. These patterns also appear in an independent corpus from the Manus Skills Marketplace.

By Hyeongjae Lee, Jihyang Cheon, Lanu Kim
arXiv AI
Aug 24

Who Delegates to AI? Evidence from 53,000 Agent Configurations

The paper introduces the Agentic Adoption Index (AAI), a new metric that captures whether workers actually delegate tasks to AI within their workflows, rather than merely measuring potential AI applicability. Using 53,000 agent skill specifications and 18,000 O*NET task statements, the authors find that occupations with high delegation differ from those previously deemed most at risk, that AAI aligns more closely with AI’s capabilities than current usage, and that adoption peaks at mid‑wage, bachelor’s‑level occupations while declining at both ends of the wage and education spectrum. The study highlights that technical availability explains much of the variation, but other factors—such as resistance to specification or professional discretion—also influence who adopts AI. whyItMatters":"The findings suggest that actual AI adoption patterns differ from prior risk assessments, indicating that factors beyond technical feasibility shape who delegates to AI, which has implications for workforce planning and policy."

By Hyeongjae Lee, Jihyang Cheon, Lanu Kim
arXiv AI
Aug 28

The Reasoning Tax: Token Economics of LLM Reasoning Across Task Types and Deployment Contexts

The paper introduces the Token Economy Score (TES), a metric that quantifies the accuracy gain of reasoning-capable large language models relative to non-reasoning baselines, normalized by token generation cost. An empirical study across 151 runs on seven diverse benchmarks shows that task structure—such as sequential inference chains—predicts higher TES, while knowledge-recall tasks yield lower TES despite difficulty. The analysis also reveals diminishing returns at higher reasoning effort and highlights how deployment context, via Reasoning Cost Share and Deployment Cost Multiplier, can alter the economic viability of reasoning workloads.

By Sachin Gopal Wani, Ajay Dholakia, David Ellison
arXiv Statistics ML
Sep 4

Low Rank for Rank: Uncertainty-Aware Task-Specific LLM Ranking under Sparse Pairwise Comparisons

The paper introduces a low‑rank framework for ranking large language models (LLMs) on task‑specific benchmarks using sparse pairwise comparisons. By modeling the task‑by‑model ability matrix as low rank, the method shares information across related tasks while preserving task‑specific differences, and it provides uncertainty‑aware ranking through debiased estimators and simultaneous confidence sets. Experiments on synthetic data and the Chatbot Arena benchmark demonstrate improved sample efficiency and tighter, better‑calibrated ranking certificates, especially in the sparse comparison regime typical of real LLM evaluations.

By Jiachun Li, David Simchi-Levi, Will Wei Sun
arXiv Computation and Language
Aug 27

SCHEDBench: A Benchmark for Evaluating LLM Constraint Faithfulness in Natural-Language Combinatorial Scheduling

SCHEDBench is a natural‑language benchmark that evaluates whether large language models (LLMs) produce schedules that remain constraint‑feasible when the same scheduling problem is expressed in different natural‑language surface forms. The benchmark covers 1,132 instances from job‑shop scheduling, resource‑constrained project scheduling, nurse rostering, and curriculum timetabling, and uses domain‑specific templates and surface‑form variations to generate varied problem statements. Experiments with thirteen frontier and open‑weight LLMs show that models are not reliably invariant to semantically equivalent renderings, with surface‑form variation reducing feasibility and increasing hard‑constraint violations, especially when constraints are reordered.

By Shrenil Shaun Sharma, Avi Sharma