arXiv AI

Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective

arXiv:2602. 02572v2 Announce Type: replace-cross Abstract: Existing alignment methods directly use the reward model learned from user preference data to optimize an LLM policy, subject to KL regularization with respect to the base policy.

arXiv AI
Sep 1

Personalized Group Relative Policy Optimization for Heterogenous Preference Alignment

The paper introduces Personalized Group Relative Policy Optimization (P‑GRPO), a new alignment framework for large language models that separates advantage estimation from immediate batch statistics. By normalizing advantages using preference‑group‑specific reward histories instead of the concurrent generation group, P‑GRPO maintains contrastive signals for distinct user preferences. Experiments across various tasks show that P‑GRPO converges faster and yields higher rewards than standard GRPO, improving alignment with heterogeneous human preferences while preserving general capabilities.

By Jialu Wang, Heinrich Peters, Asad A. Butt, Navid Hashemi, Alireza Hashemi, Pouya M. Ghari, Joseph Hoover, James Rae, Morteza Dehghani
arXiv AI
3d ago

Multi-LLM Collaborative Alignment via Stackelberg Games

The paper introduces Stackelberg Alignment, a leader‑follower framework that lets a pool of language models collaborate and improve by learning from each other’s responses. An EXP3 bandit leader adaptively selects instructions based on difficulty and discriminability, while the models act as followers, evaluating peers and learning via DPO or GRPO with Elo‑style reputation weighting and opponent matching. Experiments on diverse benchmarks show that this adaptive curriculum outperforms static baselines by up to 12‑25% and improves multi‑LLM evolution.

By Christina Hahn, Shangbin Feng, Dean Light, Swastik Roy, Hila Gonen, Yulia Tsvetkov