arXiv AI By Marcel Hussing, Liv G. d'Aliberti, Claas Voelcker, Benjamin Eysenbach, Eric Eaton

Behavior-Consistent Deep Reinforcement Learning

Read the original on arXiv AI →

The paper introduces the concept of behavior-consistent deep reinforcement learning, aiming to produce high-performing policies that remain distributionally similar across different training runs. It shows that maximum-entropy RL can control behavioral divergence by anchoring runs to a common prior, and proves that for Boltzmann policies, a temperature proportional to Q‑function disagreement limits pairwise KL divergence. Building on this, the authors propose Q‑value Expectile Disagreement (QED), a state‑dependent temperature schedule that uses double‑critic disagreement to approximate cross‑run disagreement, and demonstrate that QED reduces across‑run divergence by two orders of magnitude on 18 continuous‑control tasks without sacrificing performance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 18

Improving Online Reinforcement Learning via Bidirectional Behavior Prior Distillation

The paper introduces Bidirectional Behavior Prior Distillation (B2PD), a method that uses action‑value priors to train a conditional variational autoencoder for generating high‑value behavior support. These expert behavior priors are then distilled into the online reinforcement learning agent, reducing inefficient exploration and stabilizing policy updates. Experiments on state‑ and pixel‑based tasks show that B2PD improves sample efficiency while maintaining stable learning dynamics.

By Gong Gao, Xiao Lai, Jiaji Shen, Ning Jia, Xianhui Liu, Weidong Zhao