arXiv:2608. 12974v1 Announce Type: new Abstract: McCoy & Griffiths (2025, henceforth M&G) suggest that a Bayesian prior can be distilled into Artificial Neural Networks (ANNs) through Model-Agnostic Meta-Learning (MAML, Finn et al.
By Orr Well, Idan Tarshish, Nur Lan, Roni Katzir
arXiv:2601.17609v3 Announce Type: replace
Abstract: In domains like medicine and finance, large-scale labeled data is costly and often unavailable, leading to models trained on small datasets that st...
By Sara Rezaeimanesh, Kundan Thind, Farzan Siddiqui, Mohammad M. Ghassemi
The paper introduces a computational model that encodes symbolic knowledge as mental programs combining natural language and source code, and uses LLM-guided Bayesian learning to sequentially infer these programs. It demonstrates that this approach satisfies data‑efficiency, uncertainty handling, and flexibility, reproducing human inductive learning and active inquiry behaviors such as anchoring and garden‑pathing. In contrast, pure LLMs and classic Bayesian models either fail the task, do not match human behavior, or require prohibitive computational resources.
By Wasu Top Piriyakulkij, Sam Acquaviva, Cassidy Langenfeld, Joshua Tenenbaum, Kevin Ellis
arXiv:2510. 19698v3 Announce Type: replace Abstract: Large Language Models (LLMs) can propose rules in natural language, sidestepping the need for a predefined predicate space in traditional rule learning.
By Yang Yang, Hua XU, Zhangyi Hu, Yutao Yue
arXiv:2606. 19264v1 Announce Type: new Abstract: The knowledge encoded in large language models (LLMs) can serve as a substrate for structured reasoning over variables describing a complex world, but accessing this knowledge in a probabilistically coherent manner poses a difficult inference problem.
By Sanghyeok Choi, Henry Gouk, Esmeralda S. Whitammer
The paper presents a Bayesian framework that unifies several large‑language‑model training and evaluation paradigms—supervised fine‑tuning (SFT), few‑shot in‑context learning (ICL), and KL‑regularized reinforcement learning (RLHF/RLVR). It shows that each method can be viewed as a two‑step process: first constructing a Bayes or Gibbs posterior over outputs or actions using a prior and a utility signal, then approximating this posterior via a forward‑KL projection onto a parametric family. The authors formalize ICL and SFT as amortized weight projections, and demonstrate that reward‑weighted SFT, reward‑weighted ICL, and advantage‑weighted SFT are all special cases of forward‑KL projection onto reward‑induced posteriors, while also outlining where these equivalences hold and where they break down.
By Junxin Fan
arXiv:2407. 12288v5 Announce Type: replace-cross Abstract: The progress of machine learning over the past decade is undeniable.
By Hong Jun Jeon, Benjamin Van Roy
arXiv:2606. 09856v1 Announce Type: cross Abstract: Post-training Large Language Models (LLMs) for reasoning typically focuses on deductive tasks such as mathematics and coding where correctness is verifiable.
By Liyi Zhang, Akshay K. Jagadish, Brenden M. Lake, Thomas L. Griffiths
BayesPrompt proposes a Bayesian approach to prompt optimisation for large language models, aiming to generate prompts that are both efficient in perplexity and human readable. The authors argue that traditional optimisation methods produce unintelligible pseudoprompts due to the ill‑posed nature of the task. Their algorithm samples prompts from a posterior distribution, and experiments on real data show marked improvements over state‑of‑the‑art alternatives across several metrics.
Selecting the best response from multiple small-model samples using a stronger scorer is a simple inference-time strategy, but fails when the small model has already committed to incorrect reasoning paths. PRM guided search avoids this by scoring candidate continuations during generation, but requires a reward model trained with step-level labels.
arXiv:2609.39525v1 Announce Type: new
Abstract: Casting Bayesian inference as a neural network optimization problem targeting an amortized posterior is attractive, as it extends to otherwise intracta...
By Hans Olischl\"ager, Svenja Jedhoff, \v{S}imon Kucharsk\'y, Aayush Mishra, Stefan T. Radev, Paul B\"urkner
arXiv:2606. 01682v1 Announce Type: cross Abstract: Selecting the best response from multiple small-model samples using a stronger scorer is a simple inference-time strategy, but fails when the small model has already committed to incorrect reasoning paths.
By Atoosa Chegini, Soheil Feizi