The paper presents a Bayesian framework that unifies several large‑language‑model training and evaluation paradigms—supervised fine‑tuning (SFT), few‑shot in‑context learning (ICL), and KL‑regularized reinforcement learning (RLHF/RLVR). It shows that each method can be viewed as a two‑step process: first constructing a Bayes or Gibbs posterior over outputs or actions using a prior and a utility signal, then approximating this posterior via a forward‑KL projection onto a parametric family. The authors formalize ICL and SFT as amortized weight projections, and demonstrate that reward‑weighted SFT, reward‑weighted ICL, and advantage‑weighted SFT are all special cases of forward‑KL projection onto reward‑induced posteriors, while also outlining where these equivalences hold and where they break down.
By Junxin Fan
arXiv:2604. 06464v2 Announce Type: replace Abstract: Conformal prediction provides distribution-free prediction intervals with finite-sample coverage guarantees, and recent work by Snell \& Griffiths reframes it as Bayesian Quadrature (BQ-CP), yielding powerful data-conditional guarantees via Dirichlet posteriors over thresholds.
By Xiayin Lou, Peng Luo
arXiv:2507. 11768v3 Announce Type: replace-cross Abstract: Bayesian accounts of in-context learning face a direct objection: exact posterior predictives for exchangeable data are invariant to task-preserving order, yet transformers change next-token probabilities when the same examples are serialized differently.
By Leon Chlon, Fatima Sheaib, Zein Khamis, Maggie Chlon, Mahdi El Zein, MarcAntonio M. Awada
The paper identifies a single direction in the unembedding matrix of large language models that encodes the unigram distribution of the training corpus, acting as a Bayesian prior when the model is uncertain. By projecting the final prediction state onto this direction, the authors derive a per‑token prior loading factor, λ, which decreases as context becomes more informative and decomposes predictions into a tempered prior and a context‑driven likelihood. Experiments across four model families (Llama, Qwen, Gemma, Pythia) show that larger models rely less on the prior in high‑context settings and that manipulating λ can steer predictions toward or away from the unigram prior in KL divergence.
By Toni J. B. Liu, Jiajun Bao, Yizhou Liu, Gurbir Arora, Nicolas Boull\'e, Rapha\"el Sarfati, Christopher J. Earls
arXiv:2608. 11162v1 Announce Type: new Abstract: The Naive Bayes (NB) classifier remains a standard choice for categorical data, yet its widely used smoothing rules, such as Laplace, Lidstone, Krichevsky-Trofimov, and the $m$-estimate, all prescribe a fixed smoothing strength that ignores feature cardinality, sample size, and class imbalance, inducing a non-vanishing bias on modern high-cardinality tabular data.
By Nguyen Thai Anh, Truong Viet Vu, Tran Thien Thanh, Vo Nguyen Quoc Bao, Ngo Hoang Tu
arXiv:2607. 18804v1 Announce Type: new Abstract: In the \emph{latent posterior model} of transformer behavior, the next-token distribution arises from a posterior over latent predictive models conditioned on the context, mixed to generate continuations.
By Garrett Baker, Vinayak Pathak, Daniel Murfet, Susan Wei
arXiv:2606. 05381v1 Announce Type: new Abstract: We propose an extended family of structured spatial priors that incorporates the total variation (TV) function with $\ell_p$ norms.
By Disi Lin, Martin Berggren, Tommy L\"ofstedt
arXiv:2606. 25745v1 Announce Type: cross Abstract: Mean Field Variational Inference (MFVI) is widely understood to underestimate posterior variance.
By James Odgers, Ben Riegler, Siddharth Swaroop, Vincent Fortuin
arXiv:2609. 27976v1 Announce Type: cross Abstract: Conformal Bayes combines Bayesian posterior predictive scores with conformal calibration, but under continuous label shift both the score and calibration weight depend on the unknown response-marginal density ratio.
By Seungjin Choi
Hierarchical data is ubiquitous in the empirical sciences and is most commonly analyzed with generalized linear mixed-effects models (GLMMs). Bayesian inference for GLMMs yields calibrated uncertainty...
arXiv:2607.10912v2 Announce Type: replace
Abstract: 3D Gaussian Splatting represents scenes as finite mixtures of anisotropic Gaussians whose number of components $K$ is set by heuristic density cont...
By Aqi Dong
arXiv:2607. 02104v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as cheap, scalable judges that compare candidate outputs pairwise -- to rank responses, select models, or triage papers.
By Jian Xu, Delu Zeng, John Paisley, Qibin Zhao