← Back to all news
arXiv Machine Learning September 21, 2026 By Brenden Latham, Mehrdad Moharrami

Offline Constrained RLHF with Multiple Preference Oracles

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • reinforcement-learning
  • safety

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jun 2

Efficient Exploration for Iterative Nash Preference Optimization

arXiv:2606. 01382v1 Announce Type: cross Abstract: Preference alignment is central to improving large language models, but standard reward-based formulations can be restrictive when human preferences are cyclic, non-transitive, or otherwise not representable by a scalar reward.

By Tianlong Nan, Xiaopeng Li, Christian Kroer, Tianyi Lin
llmsfine-tuningbenchmarkssafety
More like this →
arXiv Machine Learning
Aug 4

Learning-Based Stochastic Optimal Control with Infinite-Horizon Probabilistic Constraints

arXiv:2608. 01151v1 Announce Type: cross Abstract: In this paper, we consider stochastic optimal control problems with infinite-horizon joint chance constraints.

By Francesco Cordiano, Kanghui He, Bart De Schutter
reinforcement-learning
More like this →
arXiv AI
Sep 10

Provable Pluralistic Alignment: Multi-Party RLHF under Offline Human Feedback

arXiv:2403.05006v2 Announce Type: replace-cross Abstract: Pluralistic alignment requires learning from feedback that reflects persistent and potentially conflicting stakeholder preferences while ulti...

By Huiying Zhong, Tianwei Gao, Zhiwei Steven Wu, Linjun Zhang, Weijie J. Su, Zhun Deng
reinforcement-learningsafety
More like this →
arXiv AI
Jun 17

Learning Fair Pareto-Optimal Policies in Multi-Objective Reinforcement Learning

arXiv:2606. 18111v1 Announce Type: cross Abstract: Fairness is an important aspect of decision-making in multi-objective reinforcement learning (MORL), where policies must ensure both optimality and equity across multiple, potentially conflicting objectives.

By Umer Siddique, Peilang Li, Yongcan Cao
reinforcement-learningbenchmarkssafety
More like this →
arXiv Machine Learning
Jun 8

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models

arXiv:2505. 10892v2 Announce Type: replace Abstract: Post-training LLMs with RLHF and preference optimization methods (e.

By Akhil Agnihotri, Rahul Jain, Deepak Ramachandran, Zheng Wen
llmsreinforcement-learningfine-tuningbenchmarkssafety
More like this →
arXiv Machine Learning
Sep 22

Transductive Off-policy Proximal Policy Optimization

arXiv:2406.03894v2 Announce Type: replace Abstract: Proximal Policy Optimization (PPO) is a popular model-free reinforcement learning algorithm, esteemed for its simplicity and efficacy. However, due...

By Yaozhong Gan, Renye Yan, Xiaoyang Tan, Zhe Wu, Junliang Xing
reinforcement-learning
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea