arXiv AI By Hamed Khosravi, Xiaoming Huo

Beyond "AI Helps Humans": Decision-Targeted Evaluation Design for Human-Agent Teams in the Agentic Era

Read the original on arXiv AI →

The paper introduces TEAM-Design, a rule that assigns two replay probabilities to each task—one for a human-only replay and one for an agent-only replay—based on how difficult it is to predict the missing baseline outcome and the cost of replay. It addresses the challenge of deciding whether to keep a human-AI workflow or replace it with a single actor when only one outcome can be observed after deployment. The authors prove that TEAM-Design solves the budgeted design problem and controls error rates, and demonstrate its effectiveness on clinical and coding benchmarks, noting it excels when one comparison is clearly harder than the other.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 10

Online Surrogate Repair: Decoupling High-Fidelity Feedback from Search Length in Closed-Loop Discovery

The paper introduces Online Surrogate Repair (OSR), a closed‑loop algorithm that decouples the frequency of high‑fidelity evaluations from the length of an agent’s search by selectively updating a surrogate model with sparse, high‑fidelity data. An acquisition rule determines which candidate designs receive expensive evaluations, and the resulting labels refine the surrogate for subsequent episodes. Experiments on synthetic environments and the MADE benchmark show that OSR can reduce regret more efficiently than fixed‑surrogate approaches, requiring fewer oracle queries than high‑fidelity feedback after every episode.

By Xiaotang Feng, Philip Torr, Bruno Andreis
arXiv AI
Sep 17

Designing Agentic AI Workflow Portfolios under Imperfect Selection and Compute Cost

The paper investigates how to design portfolios of agentic AI workflows that vary in reasoning strategy, verification structure, and compute cost. It proposes a portfolio-and-selector framework where multiple workflow executions are run and the best output is chosen, balancing additional compute with potential gains in accuracy. The authors develop exact and approximate optimization methods, evaluate them on three datasets, and show modest improvements over the best single workflow.

By Mojtaba Abdolmaleki, Stefanus Jasin, Boyu Wang
arXiv AI
Jun 2

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models

arXiv:2606. 01961v1 Announce Type: new Abstract: Autonomous agents are increasingly expected to support end-to-end medical-AI research workflows, moving beyond isolated prediction tasks or short-form clinical question answering.

By Junqi Liu, Salena Song, Yuhan Wang, Jiawei Mao, Hardy Chen, Xiaoke Huang, Tianhao Qi, Pengfei Guo, Yucheng Tang, Yufan He, Can Zhao, Andriy Myronenko, Dong Yang, Daguang Xu, Yuyin Zhou