arXiv:2608. 15565v1 Announce Type: new Abstract: Experience-learning agents for optimization modeling improve by storing verified skills, but existing learners admit knowledge by checking against known answers, which real ticket streams do not provide.
By Junbo Jacob Lian, Huiling Chen, Hanzhang Qin, Chung-Piaw Teo
arXiv:2606. 17529v1 Announce Type: cross Abstract: Scientific machine-learning (SciML) surrogates approximate expensive simulations, but exact expected outputs for arbitrary inputs are unavailable (the oracle problem).
By Meng Li, Xiaohua Yang, Jie Liu, Shiyu Yan
arXiv:2607. 14545v1 Announce Type: new Abstract: Machine-learned predictions can speed up offline NP-hard optimization, but asking a predictor what to do amounts to asking it to solve the problem, and committing an unchecked prediction forfeits every worst-case guarantee.
By Haifeng Li, Mo Hai
The paper investigates how AI systems that perform best‑of‑n search require different validation strategies as the search width changes. It shows that auditing only small search widths leaves a gap in reliability estimates for larger widths, and proposes retaining candidate ranks and truth labels to estimate reliability across all widths up to N. The authors derive theoretical bounds on the minimax mean‑squared error, design procedures that achieve these bounds, and demonstrate that a shared audit can significantly reduce maximum error across many widths in practical CodeRM pools.
By Ricardo Fitas
The paper introduces Online Surrogate Repair (OSR), a closed‑loop algorithm that decouples the frequency of high‑fidelity evaluations from the length of an agent’s search by selectively updating a surrogate model with sparse, high‑fidelity data. An acquisition rule determines which candidate designs receive expensive evaluations, and the resulting labels refine the surrogate for subsequent episodes. Experiments on synthetic environments and the MADE benchmark show that OSR can reduce regret more efficiently than fixed‑surrogate approaches, requiring fewer oracle queries than high‑fidelity feedback after every episode.
By Xiaotang Feng, Philip Torr, Bruno Andreis
Agentic systems have widened the gap between producing candidate outputs and reviewing them. This paper asks a practical architectural question: should domain specialization be built into an evaluator's weights, or into the rule that decides when its judgment can be trusted?