arXiv AI By Stephanie Finley, Liudas Panavas, Thomas Mikkelson, Cam Hinton, Stacey Ganss, Bradley Monton, Emily Kendall, Michelle Spradlin, Lydia Bye, Michael O'Brien, Lauren Ylvisaker, Derek Ray, Suhaas Garre, Sushant Mehta, Edwin Chen

DAYJOB: A Benchmark for Long-Horizon Professional Work

Read the original on arXiv AI →

DAYJOB is a new benchmark comprising 130 long‑horizon professional tasks created by experts in healthcare (50 tasks) and finance (80 tasks). Each task is packaged as a containerized Harbor environment and evaluated against a detailed binary rubric, requiring agents to meet every criterion to pass. In tests across 30 model configurations, the best model (Claude Opus 5.5) achieves about 24% success in both domains, while the median model scores below 3%.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 7

FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows

arXiv:2608. 06144v1 Announce Type: new Abstract: Most agent benchmarks evaluate tasks independently and cannot measure whether experience from one task helps with later tasks.

By Bo Deng (Beihang University, Qwen DianJin Team, Alibaba Cloud Computing), Kang Zhou (Qwen DianJin Team, Alibaba Cloud Computing), Lifan Guo (Qwen DianJin Team, Alibaba Cloud Computing), Chongyang Tao (Beihang University), Xuanren Chen (Beihang University), Chenggang Xie (Beihang University), Renzhao Liang (Beihang University), Feng Chen (Qwen DianJin Team, Alibaba Cloud Computing), Chi Zhang (Qwen DianJin Team, Alibaba Cloud Computing)
arXiv AI
Sep 15

KnowBench: Effort Reduction as a Unified, Deployment-Grounded Benchmark for Clinical AI

KnowBench is a new benchmark for clinical AI that measures Effort Reduction (ER), the proportion of system-generated clinical work product accepted by clinicians after expert and safety review. The metric is applied uniformly across various administrative tasks—visit notes, billing codes, orders, EHR summarization, patient summaries, and decision support—using the clinician’s review-and-attestation as ground truth. An initial deployment of Knowtex’s models achieved an aggregate ER of 97.99% across more than one million encounters in six months, with specialty-specific ER ranging from 96.8% to 98.9%.

By Jocelyn Kang, Caroline Zhang