arXiv:2607. 03453v1 Announce Type: cross Abstract: Inference-time alignment methods, such as Best-of-$N$, offer a flexible alternative to training-based alignment by using reward models to select high-quality responses generated by a reference LLM.
By Eric Lei, Hsiang Hsu, Chun-Fu Chen
arXiv:2607. 03248v1 Announce Type: cross Abstract: The alignment of large language models with human preferences is commonly achieved through Reinforcement Learning from Human Feedback or Direct Preference Optimization.
By Jialiang Wang, Xianming Liu, Xiong Zhou, Hui Liu, Haoliang Li
arXiv:2607. 02781v1 Announce Type: cross Abstract: Inference-time alignment steers a frozen language model during decoding using auxiliary reward signals, avoiding the cost of repeated weight updates.
By Yaswanth Chittepu, Ativ Joshi, Sohini Chintala, Scott Niekum
arXiv:2506. 12529v2 Announce Type: replace-cross Abstract: Preference-based Reinforcement Learning (PbRL) entails a variety of approaches for aligning models with human intent to alleviate the burden of reward engineering.
By Sara Rajaram, R. James Cotton, Fabian H. Sinz
arXiv:2606. 09635v1 Announce Type: cross Abstract: Ensuring the reliability of Large Language Models (LLMs) under distribution drift requires inference-time adaptation.
By Hankun Lin, Ruqi Zhang
arXiv:2602. 02572v2 Announce Type: replace-cross Abstract: Existing alignment methods directly use the reward model learned from user preference data to optimize an LLM policy, subject to KL regularization with respect to the base policy.
By Haichuan Wang, Tao Lin, Lingkai Kong, Ce Li, Hezi Jiang, Milind Tambe