arXiv Machine Learning By Shuangxiu (Max), Ma (Zachary), Wenhe (Zachary), Zhao

When May a Model Replace the Experiment? Audits, Licenses, and the Price of Trust in Surrogate-Driven Design

Read the original on arXiv Machine Learning →

arXiv:2608. 01378v1 Announce Type: new Abstract: Design campaigns in chemistry, materials science, and machine learning share a bottleneck: determining how good a candidate truly is requires an expensive evaluation - an experiment, a first-principles simulation, or a full training run.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 15

The geometry of AI validation: From structural blindness to reusable audits

The paper investigates how AI systems that perform best‑of‑n search require different validation strategies as the search width changes. It shows that auditing only small search widths leaves a gap in reliability estimates for larger widths, and proposes retaining candidate ranks and truth labels to estimate reliability across all widths up to N. The authors derive theoretical bounds on the minimax mean‑squared error, design procedures that achieve these bounds, and demonstrate that a shared audit can significantly reduce maximum error across many widths in practical CodeRM pools.

By Ricardo Fitas
arXiv AI
Sep 10

Online Surrogate Repair: Decoupling High-Fidelity Feedback from Search Length in Closed-Loop Discovery

The paper introduces Online Surrogate Repair (OSR), a closed‑loop algorithm that decouples the frequency of high‑fidelity evaluations from the length of an agent’s search by selectively updating a surrogate model with sparse, high‑fidelity data. An acquisition rule determines which candidate designs receive expensive evaluations, and the resulting labels refine the surrogate for subsequent episodes. Experiments on synthetic environments and the MADE benchmark show that OSR can reduce regret more efficiently than fixed‑surrogate approaches, requiring fewer oracle queries than high‑fidelity feedback after every episode.

By Xiaotang Feng, Philip Torr, Bruno Andreis