arXiv:2607. 18804v1 Announce Type: new Abstract: In the \emph{latent posterior model} of transformer behavior, the next-token distribution arises from a posterior over latent predictive models conditioned on the context, mixed to generate continuations.
By Garrett Baker, Vinayak Pathak, Daniel Murfet, Susan Wei
arXiv:2507. 11768v3 Announce Type: replace-cross Abstract: Bayesian accounts of in-context learning face a direct objection: exact posterior predictives for exchangeable data are invariant to task-preserving order, yet transformers change next-token probabilities when the same examples are serialized differently.
By Leon Chlon, Fatima Sheaib, Zein Khamis, Maggie Chlon, Mahdi El Zein, MarcAntonio M. Awada
arXiv:2606. 30440v1 Announce Type: cross Abstract: We present a complete formal proof that transformer architectures, when their internal update mechanisms satisfy a Bayes joint-distribution condition, implement exact Bayesian posterior inference.
By Haobo Yang
arXiv:2602. 04596v2 Announce Type: replace-cross Abstract: Bayes-filtered transformers are transformers meta-learned on sequences from a prior predictive distribution to approximate the corresponding posterior predictive distribution.
By Sandra Fortini, Kenyon Ng, Sonia Petrone, Judith Rousseau, Susan Wei
arXiv:2607. 22961v1 Announce Type: new Abstract: Verbalized Machine Learning (VML) parameterizes a model as a natural-language prompt that an LLM evaluates as f(x; theta).
By Yan Zhang, Shikan Lian, Shibo Li
arXiv:2606. 20538v1 Announce Type: new Abstract: Bayesian predictive inference provides a principled framework for uncertainty quantification, data efficiency, and robust generalization.
By Qingyang Zhu, Eric Karl Oermann, Kyunghyun Cho
arXiv:2511. 05963v4 Announce Type: replace Abstract: Transformers replace recurrence with a memory that grows with sequence length and self-attention that enables ad-hoc lookups over past tokens.
By Jayden Teoh, Manan Tomar, Kwangjun Ahn, Edward S. Hu, Tim Pearce, Pratyusha Sharma, Akshay Krishnamurthy, Riashat Islam, Alex Lamb, John Langford
arXiv:2607. 19379v1 Announce Type: new Abstract: Prior work has shown that transformers can perform exact Bayesian filtering within a fixed hypothesis class.
By Siddhartha R Dalal, Vishal Misra, Abhay Parekh
arXiv:2507. 01414v2 Announce Type: replace Abstract: We introduce a new family of toy problems that combine features of linear-regression-style continuous in-context learning (ICL) with discrete associative recall.
By Sultan Daniels, Dylan Davis, Dhruv Gautam, Wentinn Liao, Gireeja Ranade, Anant Sahai
arXiv:2609.17376v1 Announce Type: new
Abstract: Large language models (LLMs) trained on next-token prediction exhibit remarkable in-context learning (ICL) abilities, yet the representations that supp...
By Daniel Balcells, Andrew Jun Lee, Chirag Rastogi, Paul M. Riechers, Adam Shai, Xavier Poncini
The paper introduces the Belief Flow Filter (BFF), a generative filtering framework that encodes the evolving posterior distribution directly into flow matching model weights and updates them via test‑time gradient descent. By avoiding particle representations and Gaussian assumptions, BFF aligns structurally with Bayesian filtering and targets the recursive filtering operator. Empirical results on five physical systems—including chaotic dynamics, sparse observations, and a tokamak plasma estimation task—show that BFF outperforms existing methods in most benchmark metrics.
By Ruiqi Feng, Chongyi Wang, Tao Zhang, Tailin Wu
arXiv:2606. 15871v1 Announce Type: cross Abstract: Bayesian inference for inverse problems is run to evaluate integrals -- posterior expectations, tail probabilities, and risks -- across a stream of observations.
By Ali Siahkoohi