arXiv Machine Learning

Offline Reinforcement Learning for Distribution-Grid Protection

The paper investigates using offline reinforcement learning to improve line‑selective tripping in distribution grids. A convolutional Q‑network trained with conservative Q‑learning (CQL) processes voltage‑current phasor and impedance data, optionally with raw waveforms, to predict faulted lines. On a realistic CIGRE medium‑voltage network, the best model achieved high per‑timestep precision, recall, and F1‑score, and correctly identified the first trip action in over 98% of fault episodes, though it mis‑tripped in a notable fraction of non‑fault cases.

arXiv AI
Jun 30

Pessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning Models

arXiv:2606. 30627v1 Announce Type: cross Abstract: Conservative offline training is widely advocated as a safe foundation for subsequent online adaptation: if a policy stays close to well-supported behaviour, the argument goes, it is less likely to exploit imperfections in a learned reward model.

By Subramanyam Sahoo, Aman Chadha, Vinija Jain, Divya Chaudhary
arXiv Machine Learning
Jul 30

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?

arXiv:2607. 27203v1 Announce Type: new Abstract: Pre-training followed by fine-tuning has become the dominant recipe for learning performant policies, and in value-based reinforcement learning (RL) this raises a natural question: given a pretrained policy, should the Q-function be pretrained on offline data too?

By Perry Dong, Ron Polonsky, Dorsa Sadigh, Chelsea Fin