GridSFM is a 15‑million‑parameter physics‑inspired graph neural network that serves as a foundation model for solving AC Optimal Power Flow (AC‑OPF) across diverse grid topologies. Pretrained on 54 topologies ranging from 500 to 4,000 buses, it achieves a 2.45 % zero‑shot generation‑cost error on a held‑out 10,000‑bus case and adapts to unseen grids with only 100 solved instances using a physics‑informed fine‑tuning scheme based on Newton’s method. The authors address the disconnected feasible set of AC‑OPF by lifting and relaxing constraints with logarithmically penalized slacks, proving the resulting elastic feasible set is contractible and that solutions can be projected back onto the original feasible set.
By Luke Bhan, Weiwei Yang, Margaret Capetz, Baosen Zhang
arXiv:2410. 04818v2 Announce Type: replace-cross Abstract: We present PINCO, an unsupervised learning framework that integrates Graph Neural Networks with physics-informed neural networks for AC optimal power flow (AC-OPF) solutions.
By Anna Varbella, Damien Briens, Blazhe Gjorgiev, Giuseppe Alessio D'Inverno, Priya L. Donti, Giovanni Sansavini
The paper proposes a framework to reduce the cost of constructing scaling laws for large foundation models by treating data collection as a Bayesian optimization problem. It shows that expanding the compute budget progressively and augmenting observed configurations with surrogate-fantasized evaluations can recover a broad experimental grid, enabling accurate scaling law fitting without training every configuration. This approach can achieve computational savings of up to 10–100× compared to a full dense grid.
By Abhash Kumar Jha, Diana Alexandra Onu\c{t}u, Neeratyoy Mallik, Swagatam Haldar, Sam Laing, Niccol\`o Ajroldi, Shiwei Liu, Joaquin Vanschoren, Aaron Klein
The paper introduces a new closed‑form scaling law that extends Chinchilla’s original formula to handle data‑constrained regimes. It decomposes loss into undercapacity, undertraining, and overfitting components, saturating between an irreducible loss and an uninformed baseline. The authors validate the model on diverse architectures and domains, achieving state‑of‑the‑art RMSE across multiple LLM scaling‑law grids and enabling cost‑aware training allocations.
By Christopher M. Bryant, Hao Liu
The paper investigates why neural network surrogates sometimes improve and sometimes worsen derivative‑free optimisation performance. It identifies three key factors—role (whether the surrogate proposes candidates or replaces gradients), radius (the neighbourhood within which a local model is reliable), and room (whether the base method can still progress)—that determine when a learned local model is beneficial. Experiments on 117 benchmark instances show that providing surrogate‑approved candidates boosts success rates, while replacing gradients or ignoring the radius can reduce them.
By Chengkuo Bian, Pengcheng Xie
arXiv:2607. 07379v1 Announce Type: new Abstract: In agentic scientific machine learning (SciML), large language model (LLM) agents can discover surrogate models and select one by an automated score, typically an error metric.
By Diab W. Abueidda, Bilal Ahmed, Panos Pantidis, Mostafa E. Mobasher
arXiv:2608. 16080v1 Announce Type: new Abstract: Thermal-aware optimization of multi-die 3D integrated circuits evaluates many designs, each a costly heat-equation solve.
By Xinling Yu, Yixing Li, Ziyue Liu, Xin Ai, Zhiyu Zeng, Hai Li, Zheng Zhang
arXiv:2608. 20061v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures significantly expand model capacity without a proportional increase in computational cost.
By Nayeon Kim, Hojin Lee, Yunju Bak, Jaesun Park, Boseop Kim
arXiv:2605. 29283v2 Announce Type: replace-cross Abstract: Recent physics foundation models claim general spatiotemporal forecasting ability, yet their evaluations often collapse performance into a single average score under a fixed training distribution.
By Mengdi Chu, Yang Liu, Ayan Biswas, Han-Wei Shen
Agentic ESOpt proposes using evolution strategies (ES) instead of reinforcement learning to fine‑tune large language‑model agents for long‑horizon tasks. ES offers model scalability, flexibility, and better long‑horizon credit assignment, enabling full‑parameter optimization with minimal GPU memory. The framework samples parameter perturbations, evaluates agents with rewards, and updates online, achieving notable performance gains on WebArena‑Lite and in test‑time prompt‑parameter co‑evolution.
By Zhi Zheng, Rongsheng Chen, Yunpeng Ba, Zhenkun Wang, Yee Whye Teh, Wee Sun Lee
arXiv:2608.24479v1 Announce Type: new
Abstract: Massively parallel simulation changes the data regime in which off-policy reinforcement learning (RL) is trained, challenging stabilizers designed for...
By Zihao Wu, Hongyao Tang, Yi Ma, Huizhong Song, Pengyi Li, Yifu Yuan, Fei Ni, Jinyi Liu, Wei Wei, Jianrong Wang, Yan Zheng, Jianye Hao
The paper introduces Penalty + Sequential Linearized Feasibility Seeking (SLFS), a self‑supervised learning framework for solving multiphase AC optimal power flow (AC‑OPF) in distribution systems with topology reconfiguration. SLFS trains directly from the AC‑OPF objective and constraints using a differentiable fixed‑point power flow solver, avoiding the need for labeled optimal solutions. It achieves negligible optimality gaps and near‑zero constraint violations on IEEE feeders up to 8,500 nodes, delivering up to three orders of magnitude speedups over IPOPT while maintaining robustness to large distributional shifts.
By Hoang T. Nguyen, Shaohui Liu, Reetam Sen Biswas, Varsha Pendyala, Nurali Virani, Deepjyoti Deka, Priya L. Donti