← Back to all news
arXiv Computation and Language August 25, 2026 By Gengxu Li, Tingyu Xia, Yi Chang, Yuan Wu

Length-Controlled Margin-Based Preference Optimization without Reference Model

Read the original on arXiv Computation and Language →

The Flow has not summarised this story yet — read it at arXiv Computation and Language.

  • llms
  • reinforcement-learning
  • benchmarks
  • safety

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jun 12

Boosting Direct Preference Optimization with Penalization

arXiv:2606. 12505v1 Announce Type: cross Abstract: Offline preference optimization has become a practical substitute for reinforcement learning from human feedback, but pairwise objectives such as Direct Preference Optimization (DPO) and its variants use only the chosen and rejected responses stored in a static dataset.

By Pengwei Sun
llmsreinforcement-learning
More like this →
arXiv AI
Jun 2

Margin Adaptive DPO: Leveraging Reward Model for Granular Control in Preference Optimization

arXiv:2510. 05342v2 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO) has emerged as a simple and effective method for aligning large language models.

By Hyung Gyu Rho
llmsnlpreinforcement-learning
More like this →
arXiv AI
Jun 30

Distributionally Robust Reinforcement Learning with Human Feedback

arXiv:2503. 00539v2 Announce Type: replace-cross Abstract: Reinforcement learning from human feedback (RLHF) has evolved to be one of the main methods for fine-tuning large language models (LLMs).

By Debmalya Mandal, Paulius Sasnauskas, Goran Radanovic
llmsreinforcement-learningfine-tuning
More like this →
arXiv AI
Jun 11

Autoregressive Direct Preference Optimization

arXiv:2602. 09533v2 Announce Type: replace Abstract: Direct preference optimization (DPO) has emerged as a promising approach for aligning large language models (LLMs) with human preferences.

By Masanari Oi, Mahiro Ukai, Masahiro Kaneko, Naoaki Okazaki, Nakamasa Inoue
llmsreinforcement-learning
More like this →
arXiv Machine Learning
Jun 25

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF

arXiv:2505. 12843v2 Announce Type: replace Abstract: Reinforcement Learning from Human Feedback (RLHF) relies on reward models to align large language models with human preferences.

By Kangwen Zhao, Jianfeng Cai, Jinhua Zhu, Ruopei Sun, Dongyun Xue, Wengang Zhou, Li Li, Houqiang Li
llmsreinforcement-learningsafety
More like this →
arXiv AI
Jul 7

Adaptive Margin RLHF via Preference over Preferences

arXiv:2509. 22851v4 Announce Type: replace-cross Abstract: Margin-based optimization is fundamental to improving generalization and robustness in classification tasks.

By Yaswanth Chittepu, Prasann Singhal, Greg Durrett, Scott Niekum
reinforcement-learningsafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea