arXiv AI By Max Weltevrede, Matthijs T. J. Spaan, Wendelin B\"ohmer

Generalization in offline RL: The structure is more important than the amount of pessimism

Read the original on arXiv AI →

arXiv:2607. 02288v1 Announce Type: cross Abstract: While pessimism counteracts overestimation bias in offline reinforcement learning (RL), being overly conservative has been associated with hindering certain forms of generalization.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 18

Improving Generalization and Robustness in Offline Reinforcement Learning via Boundary-Aware Data Augmentation

The paper introduces BADA, a Boundary-Aware Data Augmentation technique for offline reinforcement learning. By interpolating neighboring states to create synthetic data that respects the original distribution, BADA improves in-distribution generalization and robustness. Experiments on limited offline datasets show that BADA achieves state-of-the-art performance across diverse benchmarks.

By Gong Gao, Weidong Zhao, Xianhui Liu