arXiv Machine Learning

Gradient-Free Topology Adaptation for Power Flow Surrogates via In-Context Whitening

arXiv:2607. 12241v1 Announce Type: cross Abstract: Machine-learned surrogates for the AC power flow (ACPF) problem amortize the cost of repeated solves on a fixed network, but lose one to two orders of magnitude of accuracy when a line outage changes the topology.

arXiv Machine Learning
Sep 25

GridSFM: A Foundation Model for Solving AC Optimal Power Flow

GridSFM is a 15‑million‑parameter physics‑inspired graph neural network that serves as a foundation model for solving AC Optimal Power Flow (AC‑OPF) across diverse grid topologies. Pretrained on 54 topologies ranging from 500 to 4,000 buses, it achieves a 2.45 % zero‑shot generation‑cost error on a held‑out 10,000‑bus case and adapts to unseen grids with only 100 solved instances using a physics‑informed fine‑tuning scheme based on Newton’s method. The authors address the disconnected feasible set of AC‑OPF by lifting and relaxing constraints with logarithmically penalized slacks, proving the resulting elastic feasible set is contractible and that solutions can be projected back onto the original feasible set.

By Luke Bhan, Weiwei Yang, Margaret Capetz, Baosen Zhang
Hugging Face Trending Papers
Jul 15

MxGPS: Multiplex Graph Transformers for a Power Grid Foundation Model

Single-task fine-tuning of graph neural networks (GNNs) for power grid problems exhibits a systematic failure mode: models that achieve the lowest in-distribution error degrade the most under topology shift. We term this topology overfitting: the tendency of task-specific gradient signals to encode relational structure particular to the training topologies rather than the underlying physics, causing models to fail on unseen grids despite strong in-distribution performance.

arXiv Machine Learning
Jul 16

Power Homotopy for Zeroth-Order Non-Convex Optimizations

arXiv:2511. 13592v2 Announce Type: replace-cross Abstract: The existing method of GS-PowerOpt solves the non-convex optimization problem of the form $\max_{\boldsymbol{x} \in \mathbb{R}^d} f(\boldsymbol{x})$ through maximizing a Gaussian-smoothed surrogate $F_{N,\sigma}(\boldsymbol{\mu}) = \mathbb{E}_{\boldsymbol{x}\sim\mathcal{N}(\boldsymbol{\mu},\sigma^2 I_d)}[e^{N f(\boldsymbol{x})}]$.

By Chen Xu
arXiv AI
Jun 2

On the Generalization in Topology Optimization via Sensitivity-Conditioned Bernoulli Flow Matching

arXiv:2606. 02179v1 Announce Type: cross Abstract: Surrogate models for topology optimization (TO) exhibit highly variable out-of-distribution (OOD) generalization under distribution shifts such as changing loads or boundary conditions, yet the source of this variability remains unclear.

By Mohammad Rashed, Duarte F. Valoroso Madeira, Babak Gholami, Caglar Guerbuez, Yunjia Yang, Nils Thuerey
arXiv AI
Jul 16

MxGPS: Multiplex Graph Transformers for a Power Grid Foundation Model

arXiv:2607. 13763v1 Announce Type: cross Abstract: Single-task fine-tuning of graph neural networks (GNNs) for power grid problems exhibits a systematic failure mode: models that achieve the lowest in-distribution error degrade the most under topology shift.

By Charilaos Papaioannou, Ioannis Tsantilas, Dimitris Giannakakos, Vasilis Michalakopoulos, Sotiris Pelekis, Vangelis Marinakis, Arsam Aryandoust, Antonello Monti, Ricardo J. Bessa, Perdo P. Vergara, Jochen Cremer, Elissaios Sarmas
arXiv Machine Learning
4d ago

Scaling Zero-Order Pretraining through Model Sharding

arXiv:2609.37899v1 Announce Type: new Abstract: Zero-order optimization (ZO) trains without backpropagation, making it relevant to forward-only hardware and non-differentiable loss, but its gradient...

By Francois Chaubard, Mykel J. Kochenderfer, Chris R\'e
arXiv Machine Learning
Sep 10

Dense Structural Compression of Transformers via Gauge-Correct Channel Removal

The paper introduces GaugeLasso, a method that applies symmetric group‑lasso penalties to transformer channels during training, enabling entire tensor slices to be zeroed out while maintaining dense tensors for GPU efficiency. By calibrating channel penalties based on inference utility per compute, the network self‑organizes into depth‑dependent structural profiles that can be dramatically smaller than the original architecture, achieving up to 255‑fold compression on a polynomial division task and outperforming hand‑designed baselines on language modeling and autoencoding benchmarks. The approach also accelerates training and reveals over‑provisioned axes that guide subsequent design iterations.

By Jed A. Duersch, Na\"im Es-Sebbani, Nathana\"el Haas, Zied Bouraoui
arXiv AI
6d ago

Hierarchical GNNs for power flow: letting physics shape the hierarchy

The paper presents a hierarchical graph neural network (GNN) approach for power‑flow modeling that leverages physics‑informed graph reductions. By exchanging information through two reduced graphs within the corrective network of GENCO, the model achieves superior generalization across three grid topologies and new operating scenarios, outperforming both the flat baseline and a Quotient construction. Training required only about 200 epochs and fewer than 1,900 scenarios per grid, yet the hierarchical models achieved a macro family‑balanced voltage error of 0.851±0.110, a 51.3% improvement over the per‑bus mean fitted on training solutions.

By Carmine Delle Femine, Leire Garin Atxaga, Asier Diaz-Iglesias, Juan Pablo Maroto Herrera, Ane Miren Florez-Tapia, Marco Quartulli, Izaro Goienetxea Urziku
arXiv Machine Learning
Sep 22

Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone

The paper introduces Neural Spectral Capacity (NSC), a closed‑form metric derived from the singular‑value spectrum of weight matrices that can be computed solely from a network’s architectural specification. Unlike traditional measures such as #Params and #FLOPs, NSC captures architectural structure (depth, width, head, FFN allocations) and can be evaluated without instantiating the model, data, or gradients. Using a dynamic‑programming solver (NSC‑DP), the authors demonstrate that NSC can efficiently identify architectures that outperform existing training‑free proxies across Transformer and CNN families, and achieve state‑of‑the‑art results in tasks such as WikiText‑103 and commonsense reasoning with LLaMA‑7B. whyItMatters":"NSC provides a fast, architecture‑only proxy that outperforms conventional metrics and training‑free proxies, enabling more effective design and pruning of large models without costly training or data."

By Chenyu Zhu, Ruoyu Zhao, Zhichao Lu