arXiv Machine Learning By Yijiang Li, Emon Dey, Stefano Fenu, Massimiliano Lupo Pasini, Teja Kuruganti, Kibaek Kim

Scaling Laws for Physics-Aware ACOPF Surrogate Learning

Read the original on arXiv Machine Learning →

The paper investigates how physics‑aware objectives, specifically the augmented Lagrangian (AL), affect the constraint satisfaction of learning‑based surrogates for AC optimal power flow (ACOPF). By sweeping model and dataset sizes under both mean‑squared error (MSE) and AL training, the authors find that both objectives improve following power‑law trends, but AL yields a slower growth in constraint violation with network size. On matched hardware, AL reduces violation by nearly 30× for an order of magnitude more training time, with negligible added memory.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 25

GridSFM: A Foundation Model for Solving AC Optimal Power Flow

GridSFM is a 15‑million‑parameter physics‑inspired graph neural network that serves as a foundation model for solving AC Optimal Power Flow (AC‑OPF) across diverse grid topologies. Pretrained on 54 topologies ranging from 500 to 4,000 buses, it achieves a 2.45 % zero‑shot generation‑cost error on a held‑out 10,000‑bus case and adapts to unseen grids with only 100 solved instances using a physics‑informed fine‑tuning scheme based on Newton’s method. The authors address the disconnected feasible set of AC‑OPF by lifting and relaxing constraints with logarithmically penalized slacks, proving the resulting elastic feasible set is contractible and that solutions can be projected back onto the original feasible set.

By Luke Bhan, Weiwei Yang, Margaret Capetz, Baosen Zhang
arXiv AI
Sep 7

Amortizing Scaling Law Construction Costs

The paper proposes a framework to reduce the cost of constructing scaling laws for large foundation models by treating data collection as a Bayesian optimization problem. It shows that expanding the compute budget progressively and augmenting observed configurations with surrogate-fantasized evaluations can recover a broad experimental grid, enabling accurate scaling law fitting without training every configuration. This approach can achieve computational savings of up to 10–100× compared to a full dense grid.

By Abhash Kumar Jha, Diana Alexandra Onu\c{t}u, Neeratyoy Mallik, Swagatam Haldar, Sam Laing, Niccol\`o Ajroldi, Shiwei Liu, Joaquin Vanschoren, Aaron Klein
arXiv Machine Learning
Sep 23

Practical Scaling Laws: Converting Compute into Performance in a Data-Constrained World

The paper introduces a new closed‑form scaling law that extends Chinchilla’s original formula to handle data‑constrained regimes. It decomposes loss into undercapacity, undertraining, and overfitting components, saturating between an irreducible loss and an uninformed baseline. The authors validate the model on diverse architectures and domains, achieving state‑of‑the‑art RMSE across multiple LLM scaling‑law grids and enabling cost‑aware training allocations.

By Christopher M. Bryant, Hao Liu
arXiv Machine Learning
Aug 27

Why and When Neural Networks Improve Local Approximation in Optimization

The paper investigates why neural network surrogates sometimes improve and sometimes worsen derivative‑free optimisation performance. It identifies three key factors—role (whether the surrogate proposes candidates or replaces gradients), radius (the neighbourhood within which a local model is reliable), and room (whether the base method can still progress)—that determine when a learned local model is beneficial. Experiments on 117 benchmark instances show that providing surrogate‑approved candidates boosts success rates, while replacing gradients or ignoring the radius can reduce them.

By Chengkuo Bian, Pengcheng Xie