arXiv AI By Alex Smolin, Bryan Wilder

Bayesian Intelligence from the Outside

Read the original on arXiv AI →

The paper introduces a Bayesian framework for assessing intelligence in agents such as language models. It shows that an agent’s reports are consistent with Bayesian intelligence if they are not fully contradictory across prompts, and it defines an intelligence order based on the informativeness of internal experiments. The work also demonstrates the challenges of aggregating coarse reports from intelligent agents, revealing that optimal aggregation can assign arbitrary weights to unexcluded states unless the agent reports a belief about the complete state of the world.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 30

Safety from Honesty in a Disinterested AI Predictor

arXiv:2606. 29657v1 Announce Type: new Abstract: As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified.

By Yoshua Bengio, Oliver Richardson, Tom\'a\v{s} Gaven\v{c}iak, Michael Cohen, Rory Svarc, Damiano Fornasiere, Gael Gendron, David Hyland, Aton Kamanda, Adam Oberman, Francis Rhys Ward, Anna Gaven\v{c}iak, Jacob Livingston Slosser, Vincent Mai, Iulian Serban, Joumana Ghosn
Hugging Face Trending Papers
Jun 28

Safety from Honesty in a Disinterested AI Predictor

As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified. We present a formal safety argument for the Scientist AI (SAI) Predictor, trained to approximate the Bayesian posterior conditioned on a dataset of "epistemically contextualized" natural-language statements.

arXiv AI
Sep 10

When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs

The paper introduces a decision‑theoretic framework that elicits both probability judgments and decisions from large language models (LLMs) to test whether their reported beliefs are consistent with their actions. It shows that this framework yields empirically testable conditions without assuming a specific utility function. In clinical diagnosis simulations, the authors find that while LLMs’ reported beliefs are not perfect reflections of the information in their decisions, the discrepancies are small for the strongest models.

By Khurram Yamin, Jingjing Tang, Santiago Cortes-Gomez, Amit Sharma, Eric Horvitz, Bryan Wilder
Hugging Face Trending Papers
Aug 18

BayesPrompt: human readable prompts that make sense

BayesPrompt proposes a Bayesian approach to prompt optimisation for large language models, aiming to generate prompts that are both efficient in perplexity and human readable. The authors argue that traditional optimisation methods produce unintelligible pseudoprompts due to the ill‑posed nature of the task. Their algorithm samples prompts from a posterior distribution, and experiments on real data show marked improvements over state‑of‑the‑art alternatives across several metrics.