arXiv Machine Learning By Gong Gao, Weidong Zhao, Xianhui Liu

Improving Generalization and Robustness in Offline Reinforcement Learning via Boundary-Aware Data Augmentation

Read the original on arXiv Machine Learning →

The paper introduces BADA, a Boundary-Aware Data Augmentation technique for offline reinforcement learning. By interpolating neighboring states to create synthetic data that respects the original distribution, BADA improves in-distribution generalization and robustness. Experiments on limited offline datasets show that BADA achieves state-of-the-art performance across diverse benchmarks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 18

Improving Online Reinforcement Learning via Bidirectional Behavior Prior Distillation

The paper introduces Bidirectional Behavior Prior Distillation (B2PD), a method that uses action‑value priors to train a conditional variational autoencoder for generating high‑value behavior support. These expert behavior priors are then distilled into the online reinforcement learning agent, reducing inefficient exploration and stabilizing policy updates. Experiments on state‑ and pixel‑based tasks show that B2PD improves sample efficiency while maintaining stable learning dynamics.

By Gong Gao, Xiao Lai, Jiaji Shen, Ning Jia, Xianhui Liu, Weidong Zhao