arXiv AI By Haichuan Wang, Tao Lin, Lingkai Kong, Ce Li, Hezi Jiang, Milind Tambe

Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective

Read the original on arXiv AI →

arXiv:2602. 02572v2 Announce Type: replace-cross Abstract: Existing alignment methods directly use the reward model learned from user preference data to optimize an LLM policy, subject to KL regularization with respect to the base policy.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.