arXiv AI

Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective

arXiv:2602. 02572v2 Announce Type: replace-cross Abstract: Existing alignment methods directly use the reward model learned from user preference data to optimize an LLM policy, subject to KL regularization with respect to the base policy.