OpenAI Blog

Confidence-Building Measures for Artificial Intelligence: Workshop proceedings

arXiv AI
Jul 9

Measuring the metacognition of AI

arXiv:2603. 29693v3 Announce Type: replace Abstract: A robust decision-making process must take into account uncertainty, especially when the choice involves inherent risks.

By Richard Servajean, Philippe Servajean
arXiv AI
Sep 2

DualStake: Dual-Path Confidence Calibration in Deep Research Agents

DualStake introduces a dual-path confidence calibration for deep research agents, adding step confidence elicitation after each retrieval step. The method shows that evidence confidence (E-Conf) after the final retrieval provides a stronger uncertainty signal than answer confidence (A-Conf), and that A-Conf is largely influenced by E-Conf. By applying margin‑clipped, confidence‑dependent stake rewards, DualStake aligns both E-Conf and A-Conf with answer correctness, improving calibration across multiple QA benchmarks without harming accuracy.

By Yinuo Xu, Yuwei Liang, Jianjie Cheng, Meng Wang, Yongcan Yu, Shuo Lu, Jian Liang
arXiv AI
Aug 13

On Benchmarking Human-Like Intelligence in Machines

arXiv:2502. 20502v2 Announce Type: replace Abstract: Recent advances in Artificial Intelligence (AI) have yielded powerful computational models that, by learning from vast amounts of human-generated data, are increasingly posited as approximate models of human cognition.

By Lance Ying, Katherine M. Collins, Lionel Wong, Ilia Sucholutsky, Ryan Liu, Adrian Weller, Tianmin Shu, Thomas L. Griffiths, Joshua B. Tenenbaum
arXiv AI
Aug 25

The Measurement Revolution? Credible Measurement and Inference in the Age of AI

The article discusses how artificial intelligence is reshaping measurement in economics by converting unstructured data into structured variables at low cost, enabling large‑scale measurement that was previously infeasible. It outlines three stages—discovery, construct definition, and observation—where AI impacts the measurement pipeline and stresses the importance of rigorous validation to ensure credible inference. The review offers guidance on navigating the shift from a single scalable measure to multiple plausible ones that can lead to differing empirical conclusions.

By Melissa Dell, Ashesh Rambachan
OpenAI Blog
Jun 21, 2016

Concrete AI safety problems

We (along with researchers from Berkeley and Stanford) are co-authors on today’s paper led by Google Brain researchers, Concrete Problems in AI Safety. The paper explores many research problems around ensuring that modern machine learning systems operate as intended.