arXiv AI

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

arXiv:2607. 25398v1 Announce Type: new Abstract: Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let it govern every action that follows.

arXiv AI
2d ago

Agents Are Systems, Not Models: Rethinking Agentic Evaluation

The paper argues that evaluating agents as fixed models is insufficient, proposing instead to treat them as configurable systems. Using a new benchmark of four scientific tasks, the authors analyze how five configuration aspects—task information, reasoning, self‑verification, time budget, and backbone model—affect performance, noting that about 54% of outcome variance arises from run‑to‑run differences even with the same settings. The study finds that providing more task information has the strongest impact, while interactions among settings (e.g., extra time only helps with adequate information or model capability) and the choice of verification tools significantly shape agent behavior.

By Luis Wiedmann, Leander Girrbach, Cordelia Schmid, Zeynep Akata
arXiv AI
3d ago

CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?

arXiv:2605.16679v3 Announce Type: replace-cross Abstract: End-to-end automation of realistic healthcare operations stresses three capabilities underrepresented in current benchmarks: policy density,...

By Haolin Chen, Deon Metelski, Leon Qi, Tao Xia, Joonyul Lee, Steve Brown, Kevin Riley, Frank Wang, T. Y. Alvin Liu, Hank Capps MD, Zeyu Tang, Xiangchen Song, Lingjing Kong, Fan Feng, Tianyi Zeng, Zhiwei Liu, Zixian Ma, Hang Jiang, Fangli Geng, Yuan Yuan, Chenyu You, Qingsong Wen, Hua Wei, Yanjie Fu, Yue Zhao, Carl Yang, Biwei Huang, Kun Zhang, Caiming Xiong, Sanmi Koyejo, Eric P. Xing, Philip S. Yu, Weiran Yao
arXiv AI
2d ago

DAYJOB: A Benchmark for Long-Horizon Professional Work

DAYJOB is a new benchmark comprising 130 long‑horizon professional tasks created by experts in healthcare (50 tasks) and finance (80 tasks). Each task is packaged as a containerized Harbor environment and evaluated against a detailed binary rubric, requiring agents to meet every criterion to pass. In tests across 30 model configurations, the best model (Claude Opus 5.5) achieves about 24% success in both domains, while the median model scores below 3%.

By Stephanie Finley, Liudas Panavas, Thomas Mikkelson, Cam Hinton, Stacey Ganss, Bradley Monton, Emily Kendall, Michelle Spradlin, Lydia Bye, Michael O'Brien, Lauren Ylvisaker, Derek Ray, Suhaas Garre, Sushant Mehta, Edwin Chen
arXiv Computation and Language
Aug 21

One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows

arXiv:2608. 19741v1 Announce Type: new Abstract: Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation, app APIs, and function calling.

By Zhuochun Li, Youngmin Ko, Ali Keramati, Nicola Ferri, Susana Palmaz Lopez Pelaez, Liang-Chun Tsai, Calvin Wang, Mirco Milletari, Tuhin Kundu, Vadim Smolyakov, Kjartan Olafsson, Tommy Guy
arXiv AI
Sep 17

PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?

The paper introduces PACT, a benchmark designed to evaluate how well enterprise AI assistants follow compliance rules when faced with various pressures such as persistent users or hurried managers. PACT covers twelve regulated domains and forty-eight realistic multi‑turn scenarios, pairing each rule with a shortcut that violates it and applying different pressures across wording and system‑prompt modes. Using PACT, the authors profile six metrics of compliance and aggregate them into a PACTScore, revealing significant variability among 22 LLM models and that even top performers misapply rules 6–10% of the time, with user pressure increasing violations by 65% on average.

By Mika Okamoto, Ansel Kaplan Erol