arXiv AI

RAMAC: Multimodal Risk-Aware Offline Reinforcement Learning and the Role of Behavior Regularization

arXiv:2510. 02695v3 Announce Type: replace-cross Abstract: In safety-critical domains where online data collection is infeasible, offline reinforcement learning (RL) is attractive only if policies achieve high returns without catastrophic lower-tail risk.

arXiv Machine Learning
Aug 28

Simple Actors and Deep Critics for Scalable Reinforcement Learning

The paper introduces LAC (Light Actor, deep Critic), an offline reinforcement learning approach that allocates model capacity to a deep critic rather than a complex actor to improve inference efficiency. It addresses three failure modes—optimization, bootstrap-noise amplification, and value-range drift—using a residual MLP backbone, n‑step bootstrap targets, and a categorical cross‑entropy loss. Experiments on OGBench show LAC matches state‑of‑the‑art diffusion and flow‑matching baselines while reducing inference latency by up to four times.

By Guhyeon Kang, Jaehwi Lee, Minhae Kwon
arXiv Machine Learning
Sep 3

DiDrive: A Risk-Aware Hierarchical Diffusion Framework for Safe Offline Reinforcement Learning in Autonomous Driving

DiDrive introduces a risk‑aware hierarchical diffusion framework for offline reinforcement learning in autonomous driving. It combines a low‑level risk‑gated encoder with a high‑level contextual modulator to filter redundant state information, and a 3DICE policy optimization that reduces out‑of‑distribution overestimation and stabilizes gradients. On the CARLA benchmark, DiDrive outperforms baselines such as IQL, CQL, and Diffusion‑QL, achieving an 85% success rate and a 4295.68 average reward in dense traffic with 60 vehicles.

By Qisong Guo, Jingtang Chen, Zhilin Chen, Pei Xu, Mingjian Fu, Wenxi Liu, Yuanlong Yu
arXiv Machine Learning
Sep 18

Improving Online Reinforcement Learning via Bidirectional Behavior Prior Distillation

The paper introduces Bidirectional Behavior Prior Distillation (B2PD), a method that uses action‑value priors to train a conditional variational autoencoder for generating high‑value behavior support. These expert behavior priors are then distilled into the online reinforcement learning agent, reducing inefficient exploration and stabilizing policy updates. Experiments on state‑ and pixel‑based tasks show that B2PD improves sample efficiency while maintaining stable learning dynamics.

By Gong Gao, Xiao Lai, Jiaji Shen, Ning Jia, Xianhui Liu, Weidong Zhao
arXiv Machine Learning
Jul 7

CDCP: Conditional Diffusion Model with Contextual Prompts for Multi-task Offline Safe Reinforcement Learning

arXiv:2607. 03903v1 Announce Type: new Abstract: Multi-task offline safe reinforcement learning (RL) promises to learn a shared optimal safe policy from offline data across multiple tasks.

By Jiayi Guan, Tianle Zhang, Li Shen, Ruiqi Zhang, Ao Zhou, Lusong Li, Guai Chen, Mengjie Li, Alois Knoll, Xiaodong He, Changjun Jiang
arXiv Machine Learning
Jul 20

Dichotomous Diffusion Policy Optimization

arXiv:2601. 00898v3 Announce Type: replace Abstract: Diffusion-based policies have gained growing popularity in solving a wide range of decision-making tasks due to their superior expressiveness and controllable generation during inference.

By Ruiming Liang, Yinan Zheng, Kexin Zheng, Tianyi Tan, Jianxiong Li, Liyuan Mao, Zhihao Wang, Guang Chen, Hangjun Ye, Jingjing Liu, Jinqiao Wang, Xianyuan Zhan
arXiv Machine Learning
Aug 31

Amortizing intractable inference in diffusion models for vision, language, and control

The paper introduces a data‑free learning objective called relative trajectory balance for training diffusion models to sample from a posterior defined by a diffusion prior and an arbitrary black‑box constraint or likelihood. It proves asymptotic correctness of this objective and demonstrates its use across vision, language, and multimodal tasks, including classifier guidance, language infilling, and text‑to‑image generation. Additionally, the method is applied to continuous control with a score‑based behavior prior, achieving state‑of‑the‑art results in offline reinforcement learning.

By Siddarth Venkatraman, Moksh Jain, Luca Scimeca, Minsu Kim, Marcin Sendera, Mohsin Hasan, Luke Rowe, Sarthak Mittal, Pablo Lemos, Emmanuel Bengio, Alexandre Adam, Jarrid Rector-Brooks, Yoshua Bengio, Glen Berseth, Esmeralda S. Whitammer