arXiv AI

The Autonomous Agency Scale: A Behavioral Framework for Measuring Self-Directed Behavior in AI Systems

arXiv:2607. 17947v1 Announce Type: new Abstract: Existing AI measurement frameworks quantify cognitive capability, task automation, or catastrophic risk, but none measure autonomous agency: the extent to which a system behaves in a self-directed way.

arXiv AI
Sep 25

Delay-of-Gratification as a Multi-Agent Survival Micro-benchmark for Long-Horizon LLMs: Social Exposure, Personas, and Tool Use Budgets

The paper introduces a new multi‑agent micro‑benchmark called Delay‑of‑Gratification, modeled after the Stanford marshmallow experiment, to evaluate large language models (LLMs) in long‑horizon, multi‑turn interactions. In the benchmark, ReAct agents use a per‑step “raise a question” tool under various constraints—social context (broadcast vs. isolated), persona traits (age, hedonic drive), and tool‑use policy (mandatory vs. optional). Across 19,200 trajectories, the study finds that most agents exhibit an early impulse to “eat,” only 75.9% persist to the end, and factors such as isolation and hedonic drive significantly influence survival and questioning behavior, with ablations showing that removing hedonic drive and age can improve completion rates.

By Olga Manakina, Igor Bogdanov, Chung-Horng Lung
arXiv AI
2d ago

Agents Are Systems, Not Models: Rethinking Agentic Evaluation

The paper argues that evaluating agents as fixed models is insufficient, proposing instead to treat them as configurable systems. Using a new benchmark of four scientific tasks, the authors analyze how five configuration aspects—task information, reasoning, self‑verification, time budget, and backbone model—affect performance, noting that about 54% of outcome variance arises from run‑to‑run differences even with the same settings. The study finds that providing more task information has the strongest impact, while interactions among settings (e.g., extra time only helps with adequate information or model capability) and the choice of verification tools significantly shape agent behavior.

By Luis Wiedmann, Leander Girrbach, Cordelia Schmid, Zeynep Akata
arXiv AI
Aug 19

On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification

The paper re‑evaluates memory‑based self‑improving agents by adding multiple runs to measure variance and by randomizing task order. It finds that agent performance is noisy in complex, multi‑step environments and that improvement depends heavily on the sequence of tasks, revealing a hidden curriculum effect. The authors suggest that underspecification of tasks and environments contributes to this fragility and demonstrate that adding detailed rubrics and feedback can partially mitigate performance drops, though gaps remain.

By Qinyuan Ye, Yu Li, Yada Pruksachatkun, Jiaxin Zhang, Chien-Sheng Wu
arXiv AI
Jul 14

AgentAbstain: Do LLM Agents Know When Not to Act?

arXiv:2607. 10059v1 Announce Type: new Abstract: Agent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain.

By Xun Liu, Yi Evie Zhang, Vira Kasprova, Parisa Rabbani, Pardis Sadat Zahraei, Tianyu Zhang, Ali Ebrahimpour-Boroojeny, Varun Chandrasekaran
arXiv AI
Sep 10

Who Delegates to AI? Evidence from Agent Configurations in Github

The paper introduces the Agentic Adoption Index (AAI), a new measure of delegated exposure that captures whether workers actually commit tasks to AI within structured workflows. Using semantic embeddings of 888,000 agent skill specifications from GitHub and 18,000 O*NET task statements, the authors find that occupations with high delegation differ from those most vulnerable to pre-AI automation, that AAI correlates more with technical capability than with current LLM use, and that for lower‑educated occupations AAI rises with wages while it falls for higher‑educated, high‑earning workers. These patterns also appear in an independent corpus from the Manus Skills Marketplace.

By Hyeongjae Lee, Jihyang Cheon, Lanu Kim