arXiv Computation and Language

ORQA: An Occupation-Realistic Question and Answer Framework for LLM Professional Knowledge

arXiv:2609. 12366v1 Announce Type: new Abstract: We present ORQA, a method for testing occupation-level knowledge in large language models.

arXiv AI
2d ago

Legal Research Bench: Measuring End-to-End Reliability in Long-Horizon Legal Research Agents

Legal Research Bench (LRB) is a new benchmark comprising 413 open-ended U.S. legal research questions, each paired with a gold answer, supporting authorities, and a binary grading rubric. The study evaluates thirteen advanced language‑model agents using web search, case‑law search, page parsing, and retrieval tools, scoring responses only when all required criteria are met and cited authorities verify. Results show that even the best model, Claude Opus 4.8, achieves full correctness on only 42.9% of questions, with performance varying by legal area and task complexity, and no clear link between more tool calls or inference cost and higher accuracy.

By Katrina Drozdov, Oliver Chen, Langston Nashold, Rayan Krishnan
arXiv AI
Sep 10

Who Delegates to AI? Evidence from Agent Configurations in Github

The paper introduces the Agentic Adoption Index (AAI), a new measure of delegated exposure that captures whether workers actually commit tasks to AI within structured workflows. Using semantic embeddings of 888,000 agent skill specifications from GitHub and 18,000 O*NET task statements, the authors find that occupations with high delegation differ from those most vulnerable to pre-AI automation, that AAI correlates more with technical capability than with current LLM use, and that for lower‑educated occupations AAI rises with wages while it falls for higher‑educated, high‑earning workers. These patterns also appear in an independent corpus from the Manus Skills Marketplace.

By Hyeongjae Lee, Jihyang Cheon, Lanu Kim
arXiv AI
Aug 25

Designing Benchmarks for Knowledge Work

The paper proposes a new way to describe benchmarks for AI systems that perform knowledge work, outlining four explicit fields: represented activity, tested setting, required work product, and evaluated result. It builds an inventory of 18 work activities from O*NET to enable activity-level reporting across occupations, and evaluates these activities for semantic coherence, algorithm sensitivity, ontology legibility, and human interpretability. The authors apply their framework to three existing benchmarks—GDPval, OfficeQA Pro, and APEX-SWE—to show how different aspects of work can be captured within the same representation.

By Yining Hua, Hongbin Na, Cyrus Ayubcha, Levi Lian
arXiv AI
6d ago

SkillFlow: Scalable and Efficient Agent Skill Retrieval System

SkillFlow is an open, multi-stage retrieval system that helps AI agents selectively load relevant skills from a large library of community-contributed SKILL.md definitions. The pipeline uses dense retrieval, two rounds of cross-encoder reranking, and LLM-based selection to balance recall and precision. Evaluations on SkillsBench and Terminal-Bench show that SkillFlow improves performance when high-quality skills are available, but retrieval alone does not help if the corpus lacks executable skills for the target domain.

By Fangzhou Li, Pagkratios Tagkopoulos, Ilias Tagkopoulos
arXiv AI
Aug 24

Who Delegates to AI? Evidence from 53,000 Agent Configurations

The paper introduces the Agentic Adoption Index (AAI), a new metric that captures whether workers actually delegate tasks to AI within their workflows, rather than merely measuring potential AI applicability. Using 53,000 agent skill specifications and 18,000 O*NET task statements, the authors find that occupations with high delegation differ from those previously deemed most at risk, that AAI aligns more closely with AI’s capabilities than current usage, and that adoption peaks at mid‑wage, bachelor’s‑level occupations while declining at both ends of the wage and education spectrum. The study highlights that technical availability explains much of the variation, but other factors—such as resistance to specification or professional discretion—also influence who adopts AI. whyItMatters":"The findings suggest that actual AI adoption patterns differ from prior risk assessments, indicating that factors beyond technical feasibility shape who delegates to AI, which has implications for workforce planning and policy."

By Hyeongjae Lee, Jihyang Cheon, Lanu Kim
arXiv AI
Jun 16

SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

arXiv:2602. 12670v4 Announce Type: replace Abstract: Agent Skills are structured packages of procedural knowledge that augment large language model (LLM) agents at inference time.

By Xiangyi Li, Yimin Liu, Wenbo Chen, Bingran You, Zonglin Di, Yifeng He, Shenghan Zheng, Kyoung Whan Choe, Jiankai Sun, Shuyi Wang, Chujun Tao, Binxu Li, Xuandong Zhao, Hejia Geng, Xiaojun Wu, Junwei Zhou, Xiaokun Chen, Hanwen Xing, Yubo Li, Qunhong Zeng, Di Wang, Yuanli Wang, Roey Ben Chaim, Penghao Jiang, Haotian Shen, Luyang Kong, Xinyi Liu, Runhui Wang, Xuanqing Liu, Jiachen Li, Xin Lan, Yueqian Lin, Wengao Ye, Junwei He, Songlin Li, Yue Zhang, Yipeng Gao, Yijiang Li, Ze Ma, Liqiang Jing, Tianyu Wang, Kaixin Li, Yiqi Xue, Haoran Lyu, Yizhuo He, Yuchen Tian, Shutong Wu, Bowei Wang, Yixuan Gao, Bo Chen, Litong Liu, Sikai Cheng, Jiajun Bao, Shuaicheng Tong, Shuwen Xu, Terry Yue Zhuo, Tinghan Ye, Qi Qi, Miao Li, Longtai Liao, Zelin Tan, Chang Shi, Xilin Tang, Srinath Tankasala, Boqin Yuan, Yaoyao Qian, Jianhong Tu, Chenguang Wang, Yizhou Sun, Wei Wang, Aaron Taylor, Ziyue Yang, Changkun Guan, Zhikang Dong, Xinyu Zhang, Steven Dillmann, Han-chung Lee, Dawn Song
arXiv AI
Jun 9

Anything2Skill: Compiling External Knowledge into Reusable Skills for Agents

arXiv:2606. 09316v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) enables agents to access external knowledge at inference time, but it primarily retrieves fragmented declarative evidence, leaving agents to repeatedly infer task procedures from passages, manuals, examples, logs, or trajectories.

By Qianjun Pan, Yutao Yang, Junsong Li, Jie Zhou, Kai Chen, Xin Li, Qin Chen, Liang He
arXiv AI
Jun 17

A Framework for Evaluating Agentic Skills at Scale

arXiv:2606. 17819v1 Announce Type: cross Abstract: Agent skills -- structured, reusable knowledge artifacts that augment LLM agent capabilities -- have been rapidly adopted in industry, yet their cross-domain impact and use across commercial and open-source models remain under-studied, and no reusable methodology exists for evaluating an individual skill.

By Maksim Shaposhnikov, Nicolas Fortuin, Simon Stipcich, Maria I. Gorinova, Amy Heineike, Rob Willoughby