arXiv Machine Learning

Elicitation without Backpropagation: Steering Model Behavior by Optimizing the Latent Posterior

arXiv:2607. 18804v1 Announce Type: new Abstract: In the \emph{latent posterior model} of transformer behavior, the next-token distribution arises from a posterior over latent predictive models conditioned on the context, mixed to generate continuations.

arXiv Machine Learning
Jun 24

LLMs are Bayesian, In Expectation, Not in Realization

arXiv:2507. 11768v3 Announce Type: replace-cross Abstract: Bayesian accounts of in-context learning face a direct objection: exact posterior predictives for exchangeable data are invariant to task-preserving order, yet transformers change next-token probabilities when the same examples are serialized differently.

By Leon Chlon, Fatima Sheaib, Zein Khamis, Maggie Chlon, Mahdi El Zein, MarcAntonio M. Awada
arXiv Machine Learning
Jun 16

Amortized mean-shift interacting particles

arXiv:2606. 15871v1 Announce Type: cross Abstract: Bayesian inference for inverse problems is run to evaluate integrals -- posterior expectations, tail probabilities, and risks -- across a stream of observations.

By Ali Siahkoohi
arXiv Machine Learning
Jul 7

Induction Heads Interpolate N-Grams

arXiv:2607. 02800v1 Announce Type: new Abstract: Induction heads are attention circuits believed to underlie in-context learning in transformers, yet a precise characterization of the estimators they implement remains elusive.

By Francesco D'Angelo, Oguz Kaan Yuksel, Swathi Shree Narashiman, Nicolas Flammarion