← Back to all news
arXiv Machine Learning September 21, 2026 By Xuening Wu, Yanlan Kang, Shenqin Yin

IncentRL: The Trade-Off Between Preference Guidance and Task Performance

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • reinforcement-learning

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Sep 1

Locally-Guided Actor-Critic: Training a Goal-conditioned Actor with a Subgoal-aware Critic

arXiv:2608.30406v1 Announce Type: new Abstract: Goal-conditioned reinforcement learning struggles with long horizons when rewards are sparse. While a planner can provide subgoals to guide a low-level...

By Olivier Serris, St\'ephane Doncieux, Olivier Sigaud
agentsreinforcement-learning
More like this →
arXiv AI
Jun 12

Boosting Direct Preference Optimization with Penalization

arXiv:2606. 12505v1 Announce Type: cross Abstract: Offline preference optimization has become a practical substitute for reinforcement learning from human feedback, but pairwise objectives such as Direct Preference Optimization (DPO) and its variants use only the chosen and rejected responses stored in a static dataset.

By Pengwei Sun
llmsreinforcement-learning
More like this →
arXiv Computation and Language
Aug 25

Length-Controlled Margin-Based Preference Optimization without Reference Model

arXiv:2502.14643v3 Announce Type: replace Abstract: Direct Preference Optimization (DPO) is a widely adopted offline algorithm for preference-based reinforcement learning from human feedback (RLHF),...

By Gengxu Li, Tingyu Xia, Yi Chang, Yuan Wu
llmsreinforcement-learningbenchmarkssafety
More like this →
arXiv AI
Aug 5

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling

arXiv:2608. 02951v1 Announce Type: cross Abstract: Preference-based reinforcement learning (PbRL) for general stochastic MDPs often requires training a reward model.

By Evan Assmus, Qining Zhang, Lei Ying
llmsreinforcement-learningroboticsfine-tuning
More like this →
arXiv AI
Jun 26

Automating Potential-based Reward Shaping with Vision Language Model Guidance

arXiv:2606. 27180v1 Announce Type: cross Abstract: Sparse rewards are inherently challenging for reinforcement learning agents as they lack intermediate feedback to guide exploration and to correctly attribute the sparse success rewards to relevant parts of the trajectory.

By Henrik M\"uller, Daniel Kudenko
llmsagentsreinforcement-learningmultimodal
More like this →
arXiv AI
Jun 2

Efficient Exploration for Iterative Nash Preference Optimization

arXiv:2606. 01382v1 Announce Type: cross Abstract: Preference alignment is central to improving large language models, but standard reward-based formulations can be restrictive when human preferences are cyclic, non-transitive, or otherwise not representable by a scalar reward.

By Tianlong Nan, Xiaopeng Li, Christian Kroer, Tianyi Lin
llmsfine-tuningbenchmarkssafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea