arXiv AI By Junxin Fan

Unifying ICL, SFT, KL-Regularized RL Through a Bayesian Lens

Read the original on arXiv AI →

The paper presents a Bayesian framework that unifies several large‑language‑model training and evaluation paradigms—supervised fine‑tuning (SFT), few‑shot in‑context learning (ICL), and KL‑regularized reinforcement learning (RLHF/RLVR). It shows that each method can be viewed as a two‑step process: first constructing a Bayes or Gibbs posterior over outputs or actions using a prior and a utility signal, then approximating this posterior via a forward‑KL projection onto a parametric family. The authors formalize ICL and SFT as amortized weight projections, and demonstrate that reward‑weighted SFT, reward‑weighted ICL, and advantage‑weighted SFT are all special cases of forward‑KL projection onto reward‑induced posteriors, while also outlining where these equivalences hold and where they break down.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jun 1

Off-the-Shelf LLMs as Process Scorers: Training-Free Alternative to PRMs for Mathematical Reasoning

Selecting the best response from multiple small-model samples using a stronger scorer is a simple inference-time strategy, but fails when the small model has already committed to incorrect reasoning paths. PRM guided search avoids this by scoring candidate continuations during generation, but requires a reward model trained with step-level labels.