← Back to all news
arXiv Computation and Language October 1, 2026 By Khoi Le, Tri Cao, Phong Nguyen, Cong-Duy Nguyen, Anh Tuan Luu, Miao Chunyan, See-Kiong Ng, Thong Nguyen

When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift

Read the original on arXiv Computation and Language →

The Flow has not summarised this story yet — read it at arXiv Computation and Language.

  • reinforcement-learning
  • fine-tuning

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Computation and Language
Sep 1

Low-Resource Preference Adaptation of LLMs via Activation-Based Label Propagation

arXiv:2608.30902v1 Announce Type: new Abstract: Adapting large language models to user-specific preferences is often constrained by the cost of human annotation, making preference optimisation imprac...

By Alessio Galatolo, Meriem Beloucif
llmssafety
More like this →
arXiv AI
Jul 7

Adaptive Margin RLHF via Preference over Preferences

arXiv:2509. 22851v4 Announce Type: replace-cross Abstract: Margin-based optimization is fundamental to improving generalization and robustness in classification tasks.

By Yaswanth Chittepu, Prasann Singhal, Greg Durrett, Scott Niekum
reinforcement-learningsafety
More like this →
arXiv Machine Learning
2d ago

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders

arXiv:2605.16339v2 Announce Type: replace Abstract: Preference learning in large language models relies on reward models as proxies for human judgment. However, these models frequently exhibit prefer...

By Shunchang Liu, Xin Chen, Belen Martin Urcelay, Francesco Croce
llmsdiffusionreinforcement-learningbenchmarkssafety
More like this →
arXiv AI
Jul 21

Normalized Rewards for Preference Optimization

arXiv:2607. 16240v1 Announce Type: cross Abstract: Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs with human preferences.

By Shawn Im, Federico Danieli, Skyler Seto, Barry-John Theobald, Katherine Metcalf
llmsreinforcement-learningbenchmarkssafety
More like this →
arXiv AI
Jun 30

Distributionally Robust Reinforcement Learning with Human Feedback

arXiv:2503. 00539v2 Announce Type: replace-cross Abstract: Reinforcement learning from human feedback (RLHF) has evolved to be one of the main methods for fine-tuning large language models (LLMs).

By Debmalya Mandal, Paulius Sasnauskas, Goran Radanovic
llmsreinforcement-learningfine-tuning
More like this →
arXiv AI
Jun 2

Margin Adaptive DPO: Leveraging Reward Model for Granular Control in Preference Optimization

arXiv:2510. 05342v2 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO) has emerged as a simple and effective method for aligning large language models.

By Hyung Gyu Rho
llmsnlpreinforcement-learning
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea