arXiv:2606. 19353v1 Announce Type: cross Abstract: In-Context Learning (ICL) allows LLMs to adapt to new tasks from a few demonstrations, but its reliability remains a concern: predictions are highly sensitive to both prompt design and the model's ability to understand the context, obscuring whether failures arise from data properties or model limitations.
By Jinseok Chung, Minkyoung Song, Hyunji Jung, Namhoon Lee
arXiv:2604. 23099v2 Announce Type: replace-cross Abstract: Evaluating generative AI models is increasingly resource-intensive due to slow inference, expensive raters, and a rapidly growing landscape of models and benchmarks.
By Yizheng Huang, Wenjun Zeng, Aditi Kumaresan, Zi Wang
arXiv:2607. 00972v1 Announce Type: new Abstract: Trustworthy deployment of Agentic Retrieval-Augmented Generation (RAG) systems requires mechanisms for estimating when multi-stage reasoning pipelines may fail.
By Louis Donaldson, Connor Walker, Koorosh Aslansefat, Yiannis Papadopoulos
The paper investigates how Large Language Models can be used to approximate domain expert priors for Bayesian Networks by extracting probabilistic knowledge about real‑world events. Experiments on eighty publicly available networks across domains such as healthcare and finance show that LLM‑derived conditional probabilities outperform random, uniform, and next‑token baselines. The authors also demonstrate that these LLM‑generated priors can refine data‑driven distributions, especially when data is scarce, and provide the first comprehensive baseline for evaluating LLM performance in probabilistic knowledge extraction.
By Aliakbar Nafar, Kristen Brent Venable, Zijun Cui, Parisa Kordjamshidi
arXiv:2606. 00680v1 Announce Type: new Abstract: Offline reinforcement learning (RL) aims to optimize policies from pre-collected datasets.
By Hongqiang Lin, Pengfei Wang, Nenggan Zheng
arXiv:2606. 10777v1 Announce Type: new Abstract: Uncertainty estimation is critical for deploying machine learning models in high-stakes settings.
By Arthur Hoarau
arXiv:2608. 16564v1 Announce Type: new Abstract: Machine learning (ML) is a key technology driving innovation today, but ensuring ML safety remains a major challenge for safety-related applications.
By Benjamin Herd, Jessica Kelly, Mario Trapp
Machine learning (ML) is a key technology driving innovation today, but ensuring ML safety remains a major challenge for safety-related applications. A promising idea is to build proven-in-use arguments from field data, e.
arXiv:2606. 29681v1 Announce Type: new Abstract: Probabilistic model checking for Markov decision processes (MDPs) provides quantitative guarantees, but often offers limited insight into why undesired outcomes occur.
By Ryohei Oura, Georgios Fainekos, Hideki Okamoto, Bardh Hoxha
arXiv:2602. 18266v2 Announce Type: replace Abstract: Automated methods for discovering mechanistic simulator models from observational data offer a promising path toward accelerating scientific progress.
By Stefan Wahl, Raphaela Schenk, Ali Farnoud, Jakob H. Macke, Daniel Gedon
arXiv:2609.24419v1 Announce Type: cross
Abstract: Current experimental scientists increasingly rely on simulation-based inference (SBI) to invert complex models with intractable likelihoods. A primar...
By Luben M. C. Cabezas, Pedro L. C. Rodrigues, Rafael Izbicki
arXiv:2609.09855v1 Announce Type: new
Abstract: Although probabilistic statements are ubiquitous, foundational disagreements persist about their understanding, as exemplified by debates between Bayes...
By Benedikt H\"oltgen