arXiv Machine Learning By Julian Oelhaf, Alexander Luce, Christian Bergler, Andreas Maier, Siming Bayer

Offline Reinforcement Learning for Distribution-Grid Protection

Read the original on arXiv Machine Learning →

The paper investigates using offline reinforcement learning to improve line‑selective tripping in distribution grids. A convolutional Q‑network trained with conservative Q‑learning (CQL) processes voltage‑current phasor and impedance data, optionally with raw waveforms, to predict faulted lines. On a realistic CIGRE medium‑voltage network, the best model achieved high per‑timestep precision, recall, and F1‑score, and correctly identified the first trip action in over 98% of fault episodes, though it mis‑tripped in a notable fraction of non‑fault cases.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 30

Pessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning Models

arXiv:2606. 30627v1 Announce Type: cross Abstract: Conservative offline training is widely advocated as a safe foundation for subsequent online adaptation: if a policy stays close to well-supported behaviour, the argument goes, it is less likely to exploit imperfections in a learned reward model.

By Subramanyam Sahoo, Aman Chadha, Vinija Jain, Divya Chaudhary