arXiv Machine Learning By Jiayi Guan, Tianle Zhang, Li Shen, Ruiqi Zhang, Ao Zhou, Lusong Li, Guai Chen, Mengjie Li, Alois Knoll, Xiaodong He, Changjun Jiang

CDCP: Conditional Diffusion Model with Contextual Prompts for Multi-task Offline Safe Reinforcement Learning

Read the original on arXiv Machine Learning →

arXiv:2607. 03903v1 Announce Type: new Abstract: Multi-task offline safe reinforcement learning (RL) promises to learn a shared optimal safe policy from offline data across multiple tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
6d ago

Q-learning Penalized Transformer for Safe Offline Reinforcement Learning

The paper introduces Q-learning Penalized Transformer (QPT), a training–inference consistent framework for safe offline reinforcement learning. QPT trains a Transformer policy that generates actions conditioned on trajectory context and target return/cost while incorporating a Q-shaped penalty to balance safety, reward maximization, and behavior regularization. The method consistently outperforms strong baselines on 38 DSRL benchmark tasks and adapts robustly to varying constraint thresholds.

arXiv AI
Aug 19

Task Specialization Fine-Tuning for Contextual Reinforcement Learning

The paper introduces Task Specialization Fine-Tuning (TSFT), an online framework that allocates a limited fine‑tuning budget across multiple task regions in Contextual Reinforcement Learning. TSFT predicts fine‑tuning performance with a simple parametric model and solves the budget allocation problem exactly using integer linear programming. Experiments on combinatorial optimization, continuous control, and LLM fine‑tuning show that TSFT outperforms baselines in task coverage and approaches oracle performance.

By Jianan Zhou, Jung-Hoon Cho, Tianyue Zhou, Han Zheng, Jie Zhang, Roy Dong, Yining Ma, Cathy Wu
arXiv Machine Learning
Jul 27

Safe In-Context Reinforcement Learning

arXiv:2509. 25582v4 Announce Type: replace Abstract: In-context reinforcement learning (ICRL) is an emerging RL paradigm where an agent, after pretraining, can adapt to out-of-distribution test tasks without any parameter updates, instead relying on an expanding context of interaction history.

By Amir Moeini, Minjae Kwon, Alper Kamil Bozkurt, Yuichi Motai, Rohan Chandra, Lu Feng, Shangtong Zhang
arXiv AI
Sep 21

OneBid: A Unified Auto-Bidding Foundation Model for Diverse oCPX Advertising Scenarios

OneBid is a unified auto‑bidding foundation model that consolidates diverse cost‑per‑X (oCPX) advertising scenarios into a single framework. It builds on Decision Transformer by conditioning on two atomic signals—Return‑to‑Go for conversion value and Cost‑to‑Go for cost ratio—and incorporates value‑aware regularization. A sequence‑level Mixture‑of‑Experts architecture captures cross‑scenario knowledge while preserving low latency, and a Critic‑guided Relative Offline Policy optimization (CROP) aligns the backbone with scenario‑specific preferences without unsafe online exploration. In production at Kuaishou, OneBid achieved a 2.2% overall ADVV increase and up to 13.1% in the ROAS scenario.

By Yewen Li, Peng Jiang, Yitian Li, Pengfei Lv, Xialong Liu, Peng Jiang, Qingpeng Cai
arXiv Machine Learning
Jul 20

Dichotomous Diffusion Policy Optimization

arXiv:2601. 00898v3 Announce Type: replace Abstract: Diffusion-based policies have gained growing popularity in solving a wide range of decision-making tasks due to their superior expressiveness and controllable generation during inference.

By Ruiming Liang, Yinan Zheng, Kexin Zheng, Tianyi Tan, Jianxiong Li, Liyuan Mao, Zhihao Wang, Guang Chen, Hangjun Ye, Jingjing Liu, Jinqiao Wang, Xianyuan Zhan