GridSFM is a 15‑million‑parameter physics‑inspired graph neural network that serves as a foundation model for solving AC Optimal Power Flow (AC‑OPF) across diverse grid topologies. Pretrained on 54 topologies ranging from 500 to 4,000 buses, it achieves a 2.45 % zero‑shot generation‑cost error on a held‑out 10,000‑bus case and adapts to unseen grids with only 100 solved instances using a physics‑informed fine‑tuning scheme based on Newton’s method. The authors address the disconnected feasible set of AC‑OPF by lifting and relaxing constraints with logarithmically penalized slacks, proving the resulting elastic feasible set is contractible and that solutions can be projected back onto the original feasible set.
By Luke Bhan, Weiwei Yang, Margaret Capetz, Baosen Zhang
arXiv:2410. 04818v2 Announce Type: replace-cross Abstract: We present PINCO, an unsupervised learning framework that integrates Graph Neural Networks with physics-informed neural networks for AC optimal power flow (AC-OPF) solutions.
By Anna Varbella, Damien Briens, Blazhe Gjorgiev, Giuseppe Alessio D'Inverno, Priya L. Donti, Giovanni Sansavini
The paper proposes a framework to reduce the cost of constructing scaling laws for large foundation models by treating data collection as a Bayesian optimization problem. It shows that expanding the compute budget progressively and augmenting observed configurations with surrogate-fantasized evaluations can recover a broad experimental grid, enabling accurate scaling law fitting without training every configuration. This approach can achieve computational savings of up to 10–100× compared to a full dense grid.
By Abhash Kumar Jha, Diana Alexandra Onu\c{t}u, Neeratyoy Mallik, Swagatam Haldar, Sam Laing, Niccol\`o Ajroldi, Shiwei Liu, Joaquin Vanschoren, Aaron Klein
The paper introduces a new closed‑form scaling law that extends Chinchilla’s original formula to handle data‑constrained regimes. It decomposes loss into undercapacity, undertraining, and overfitting components, saturating between an irreducible loss and an uninformed baseline. The authors validate the model on diverse architectures and domains, achieving state‑of‑the‑art RMSE across multiple LLM scaling‑law grids and enabling cost‑aware training allocations.
By Christopher M. Bryant, Hao Liu
The paper investigates why neural network surrogates sometimes improve and sometimes worsen derivative‑free optimisation performance. It identifies three key factors—role (whether the surrogate proposes candidates or replaces gradients), radius (the neighbourhood within which a local model is reliable), and room (whether the base method can still progress)—that determine when a learned local model is beneficial. Experiments on 117 benchmark instances show that providing surrogate‑approved candidates boosts success rates, while replacing gradients or ignoring the radius can reduce them.
By Chengkuo Bian, Pengcheng Xie
arXiv:2607. 07379v1 Announce Type: new Abstract: In agentic scientific machine learning (SciML), large language model (LLM) agents can discover surrogate models and select one by an automated score, typically an error metric.
By Diab W. Abueidda, Bilal Ahmed, Panos Pantidis, Mostafa E. Mobasher