arXiv:2607. 12338v1 Announce Type: new Abstract: Agent benchmarks often compare two agents after all tasks have run, but costly evaluations make partial runs tempting.
By Wei-Jung Huang
The paper introduces Online Surrogate Repair (OSR), a closed‑loop algorithm that decouples the frequency of high‑fidelity evaluations from the length of an agent’s search by selectively updating a surrogate model with sparse, high‑fidelity data. An acquisition rule determines which candidate designs receive expensive evaluations, and the resulting labels refine the surrogate for subsequent episodes. Experiments on synthetic environments and the MADE benchmark show that OSR can reduce regret more efficiently than fixed‑surrogate approaches, requiring fewer oracle queries than high‑fidelity feedback after every episode.
By Xiaotang Feng, Philip Torr, Bruno Andreis
Repeated evaluation can estimate a benchmark score accurately while still requiring replication to certify narrow uncertainty. We characterize that requirement on a fixed grid of $M$ tasks with $L$ binary paths per task under the hard budget $(M+t)K$, where each path costs at most $K$ responses or episodes.
The paper investigates how to design portfolios of agentic AI workflows that vary in reasoning strategy, verification structure, and compute cost. It proposes a portfolio-and-selector framework where multiple workflow executions are run and the best output is chosen, balancing additional compute with potential gains in accuracy. The authors develop exact and approximate optimization methods, evaluate them on three datasets, and show modest improvements over the best single workflow.
By Mojtaba Abdolmaleki, Stefanus Jasin, Boyu Wang
arXiv:2609. 29140v1 Announce Type: new Abstract: Repeated evaluation can estimate a benchmark score accurately while still requiring replication to certify narrow uncertainty.
By Yezhou Cheng, Runjia Du, Zeming Liu, Qibai Chen, Hang Lyu, Yilan Wei, Yankai Zeng, Bojun Lin
arXiv:2606. 01961v1 Announce Type: new Abstract: Autonomous agents are increasingly expected to support end-to-end medical-AI research workflows, moving beyond isolated prediction tasks or short-form clinical question answering.
By Junqi Liu, Salena Song, Yuhan Wang, Jiawei Mao, Hardy Chen, Xiaoke Huang, Tianhao Qi, Pengfei Guo, Yucheng Tang, Yufan He, Can Zhao, Andriy Myronenko, Dong Yang, Daguang Xu, Yuyin Zhou