arXiv:2310.04363v3 Announce Type: replace
Abstract: Autoregressive large language models (LLMs) compress knowledge from their training data through next-token conditional distributions. This limits t...
By Edward J. Hu, Moksh Jain, Eric Elmoznino, Younesse Kaddar, Guillaume Lajoie, Yoshua Bengio, Esmeralda S. Whitammer
arXiv:2609.15007v1 Announce Type: cross
Abstract: Large language models are increasingly used as natural-language interfaces to structured data, yet they remain unreliable when answers require consis...
By Jackson Hassell, Chen Shen, Estevam Hruschka
arXiv:2605. 18476v2 Announce Type: replace-cross Abstract: Coding and computation remain major bottlenecks in Markov chain Monte Carlo (MCMC) workflows, especially as modern sampling algorithms have become increasingly complex and existing probabilistic programming systems remain limited in model support, extensibility, and composability.
By Jungang Zou, Alex Ziyu Jiang, Qixuan Chen
arXiv:2606. 09856v1 Announce Type: cross Abstract: Post-training Large Language Models (LLMs) for reasoning typically focuses on deductive tasks such as mathematics and coding where correctness is verifiable.
By Liyi Zhang, Akshay K. Jagadish, Brenden M. Lake, Thomas L. Griffiths
The paper introduces a flexible, theoretically grounded framework for steering and scaling autoregressive large language models (LLMs) through sampling. It presents two algorithms—Sequential Monte Carlo (SMC) and Replica Exchange (RE)—that guide generation toward desired distributions such as powering, product, or tilting of the base model. Experiments show these methods outperform Best‑of‑N and standard MCMC baselines, offering a systematic recipe for probabilistic inference with LLMs via sampling.
By Jiajun He, Zongyu Guo, Jos\'e Miguel Hern\'andez-Lobato, Yuanqi Du
The paper investigates how Large Language Models can be used to approximate domain expert priors for Bayesian Networks by extracting probabilistic knowledge about real‑world events. Experiments on eighty publicly available networks across domains such as healthcare and finance show that LLM‑derived conditional probabilities outperform random, uniform, and next‑token baselines. The authors also demonstrate that these LLM‑generated priors can refine data‑driven distributions, especially when data is scarce, and provide the first comprehensive baseline for evaluating LLM performance in probabilistic knowledge extraction.
By Aliakbar Nafar, Kristen Brent Venable, Zijun Cui, Parisa Kordjamshidi
arXiv:2606. 14943v1 Announce Type: cross Abstract: Causal Transformers model sequences through an autoregressive factorization of the joint distribution, which enables efficient left-to-right decoding and conditional likelihood computation.
By Yinhan Lu, Eric Elmoznino, L\'eo Gagnon, Sarthak Mittal, Tejas Kasetty, Guillaume Lajoie
arXiv:2607. 22961v1 Announce Type: new Abstract: Verbalized Machine Learning (VML) parameterizes a model as a natural-language prompt that an LLM evaluates as f(x; theta).
By Yan Zhang, Shikan Lian, Shibo Li
arXiv:2602.02427v3 Announce Type: replace
Abstract: Large Language Models (LLMs) have achieved significant breakthroughs across various domains, but they can still produce unreliable or misleading ou...
By Qihao Wen, Jiahao Wang, Yang Nan, Pengfei He, Ravi Tandon, Han Xu
arXiv:2509. 21474v4 Announce Type: replace Abstract: While diffusion language models (DLMs) have achieved competitive performance in text generation, improving their reasoning ability with reinforcement learning remains an active research area.
By Guanghan Wang, Gilad Turok, Yair Schiff, Marianne Arriola, Volodymyr Kuleshov
The paper proposes a fragment‑based reasoning framework for large language model–based machine translation. It extracts parallel source‑target fragments from retrieved similar examples and uses these fragments as intermediate reasoning traces to generate the final translation. Experiments with the Qwen3 model across six languages and multiple domains show that this approach outperforms standard k‑shot or basic drafting methods.
By Maxime Bouthors, Josep Crego, Fran\c{c}ois Yvon
arXiv:2609.36788v1 Announce Type: new
Abstract: Incorporating rich task-relevant context, such as domain knowledge and external observations, is a key capability yet remains challenging for Bayesian...
By Zhongwei Yu, Sourabh Roy, Bin Cao, Xue Yan, Anjie Liu, Jun Wang