arXiv Machine Learning

PAC--Bayes Bounds on Quotient Parameter Spaces: Geometry-induced Implicit-Bias Priors

arXiv:2607. 18422v1 Announce Type: new Abstract: Overparameterized models often have continuous parameter symmetries, so different parameters define the same predictor.

arXiv AI
Sep 7

Unifying ICL, SFT, KL-Regularized RL Through a Bayesian Lens

The paper presents a Bayesian framework that unifies several large‑language‑model training and evaluation paradigms—supervised fine‑tuning (SFT), few‑shot in‑context learning (ICL), and KL‑regularized reinforcement learning (RLHF/RLVR). It shows that each method can be viewed as a two‑step process: first constructing a Bayes or Gibbs posterior over outputs or actions using a prior and a utility signal, then approximating this posterior via a forward‑KL projection onto a parametric family. The authors formalize ICL and SFT as amortized weight projections, and demonstrate that reward‑weighted SFT, reward‑weighted ICL, and advantage‑weighted SFT are all special cases of forward‑KL projection onto reward‑induced posteriors, while also outlining where these equivalences hold and where they break down.

By Junxin Fan
arXiv Machine Learning
Jun 19

Weighted Bayesian Conformal Prediction

arXiv:2604. 06464v2 Announce Type: replace Abstract: Conformal prediction provides distribution-free prediction intervals with finite-sample coverage guarantees, and recent work by Snell \& Griffiths reframes it as Bayesian Quadrature (BQ-CP), yielding powerful data-conditional guarantees via Dirichlet posteriors over thresholds.

By Xiayin Lou, Peng Luo
arXiv Machine Learning
Jun 24

LLMs are Bayesian, In Expectation, Not in Realization

arXiv:2507. 11768v3 Announce Type: replace-cross Abstract: Bayesian accounts of in-context learning face a direct objection: exact posterior predictives for exchangeable data are invariant to task-preserving order, yet transformers change next-token probabilities when the same examples are serialized differently.

By Leon Chlon, Fatima Sheaib, Zein Khamis, Maggie Chlon, Mahdi El Zein, MarcAntonio M. Awada
arXiv AI
Sep 4

The Geometry of Ignorance: LLMs Know When to Temper Bayesian Priors

The paper identifies a single direction in the unembedding matrix of large language models that encodes the unigram distribution of the training corpus, acting as a Bayesian prior when the model is uncertain. By projecting the final prediction state onto this direction, the authors derive a per‑token prior loading factor, λ, which decreases as context becomes more informative and decomposes predictions into a tempered prior and a context‑driven likelihood. Experiments across four model families (Llama, Qwen, Gemma, Pythia) show that larger models rely less on the prior in high‑context settings and that manipulating λ can steer predictions toward or away from the unigram prior in KL divergence.

By Toni J. B. Liu, Jiajun Bao, Yizhou Liu, Gurbir Arora, Nicolas Boull\'e, Rapha\"el Sarfati, Christopher J. Earls
arXiv Machine Learning
Aug 12

Hierarchical Empirical-Bayes Naive Bayes: Minimax Smoothing and Calibration with AODE Extension

arXiv:2608. 11162v1 Announce Type: new Abstract: The Naive Bayes (NB) classifier remains a standard choice for categorical data, yet its widely used smoothing rules, such as Laplace, Lidstone, Krichevsky-Trofimov, and the $m$-estimate, all prescribe a fixed smoothing strength that ignores feature cardinality, sample size, and class imbalance, inducing a non-vanishing bias on modern high-cardinality tabular data.

By Nguyen Thai Anh, Truong Viet Vu, Tran Thien Thanh, Vo Nguyen Quoc Bao, Ngo Hoang Tu