arXiv:2609.34426v2 Announce Type: replace
Abstract: This paper addresses the problem of safe offline reinforcement learning, which involves training a policy to satisfy safety constraints using an of...
By Shengchao Hu, Peng Wang, Jifeng Hu, Qiyang Zhou, Anning Hu, Li Shen, Ya Zhang, Dacheng Tao
The paper introduces Q-learning Penalized Transformer (QPT), a training–inference consistent framework for safe offline reinforcement learning. QPT trains a Transformer policy that generates actions conditioned on trajectory context and target return/cost while incorporating a Q-shaped penalty to balance safety, reward maximization, and behavior regularization. The method consistently outperforms strong baselines on 38 DSRL benchmark tasks and adapts robustly to varying constraint thresholds.
The paper introduces Task Specialization Fine-Tuning (TSFT), an online framework that allocates a limited fine‑tuning budget across multiple task regions in Contextual Reinforcement Learning. TSFT predicts fine‑tuning performance with a simple parametric model and solves the budget allocation problem exactly using integer linear programming. Experiments on combinatorial optimization, continuous control, and LLM fine‑tuning show that TSFT outperforms baselines in task coverage and approaches oracle performance.
By Jianan Zhou, Jung-Hoon Cho, Tianyue Zhou, Han Zheng, Jie Zhang, Roy Dong, Yining Ma, Cathy Wu
arXiv:2509. 25582v4 Announce Type: replace Abstract: In-context reinforcement learning (ICRL) is an emerging RL paradigm where an agent, after pretraining, can adapt to out-of-distribution test tasks without any parameter updates, instead relying on an expanding context of interaction history.
By Amir Moeini, Minjae Kwon, Alper Kamil Bozkurt, Yuichi Motai, Rohan Chandra, Lu Feng, Shangtong Zhang
OneBid is a unified auto‑bidding foundation model that consolidates diverse cost‑per‑X (oCPX) advertising scenarios into a single framework. It builds on Decision Transformer by conditioning on two atomic signals—Return‑to‑Go for conversion value and Cost‑to‑Go for cost ratio—and incorporates value‑aware regularization. A sequence‑level Mixture‑of‑Experts architecture captures cross‑scenario knowledge while preserving low latency, and a Critic‑guided Relative Offline Policy optimization (CROP) aligns the backbone with scenario‑specific preferences without unsafe online exploration. In production at Kuaishou, OneBid achieved a 2.2% overall ADVV increase and up to 13.1% in the ROAS scenario.
By Yewen Li, Peng Jiang, Yitian Li, Pengfei Lv, Xialong Liu, Peng Jiang, Qingpeng Cai
arXiv:2601. 00898v3 Announce Type: replace Abstract: Diffusion-based policies have gained growing popularity in solving a wide range of decision-making tasks due to their superior expressiveness and controllable generation during inference.
By Ruiming Liang, Yinan Zheng, Kexin Zheng, Tianyi Tan, Jianxiong Li, Liyuan Mao, Zhihao Wang, Guang Chen, Hangjun Ye, Jingjing Liu, Jinqiao Wang, Xianyuan Zhan