OpenAI Blog

Measuring Goodhart’s law

Goodhart’s law famously says: “When a measure becomes a target, it ceases to be a good measure. ” Although originally from economics, it’s something we have to grapple with at OpenAI when figuring out how to optimize objectives that are difficult or costly to measure.

arXiv Machine Learning
Jul 31

Procedural Fairness in Multi-Agent Bandits

arXiv:2601. 10600v2 Announce Type: replace-cross Abstract: In the context of multi-agent multi-armed bandits (MA-MAB), fairness is often reduced to outcomes: maximizing welfare, reducing inequality, or balancing utilities.

By Joshua Caiata, Carter Blair, Kate Larson
arXiv AI
Aug 25

The Measurement Revolution? Credible Measurement and Inference in the Age of AI

The article discusses how artificial intelligence is reshaping measurement in economics by converting unstructured data into structured variables at low cost, enabling large‑scale measurement that was previously infeasible. It outlines three stages—discovery, construct definition, and observation—where AI impacts the measurement pipeline and stresses the importance of rigorous validation to ensure credible inference. The review offers guidance on navigating the shift from a single scalable measure to multiple plausible ones that can lead to differing empirical conclusions.

By Melissa Dell, Ashesh Rambachan
arXiv Machine Learning
Jun 9

Proper Calibeating

arXiv:2605. 26703v2 Announce Type: replace-cross Abstract: The classic concept of "calibrated forecasts" and its more recent refinement, "calibeating," are defined with respect to the standard quadratic scoring rule.

By Dean P. Foster, Sergiu Hart
arXiv Computation and Language
Sep 1

Bye Bye Perspective API: Lessons for Building and Governing Measurement Infrastructure

arXiv:2604.25580v2 Announce Type: replace Abstract: Perspective API closes at the end of 2026, removing the de facto standard for toxicity measurement and exposing researchers' dependence on a tool t...

By David Hartmann, Manuel Tonneau, Angelie Kraft, LK Seiling, Dimitri Staufer, Pieter Delobelle, Jan Fillies, Anna Ricarda Luther, Jan Batzner, Mareike Lisker
arXiv Machine Learning
Sep 18

When fairness metrics fail: A utility-based perspective on $\varepsilon$-fairness

The paper argues that traditional probabilistic fairness metrics can miss significant disparities in the actual consequences of decisions. By introducing a utility-based framework, the authors show that a process can satisfy ε-fairness yet still be maximally unfair when utilities are considered. They apply this framework to college admissions and credit‑risk assessment, demonstrating that equalizing probabilities alone may mask unequal utility outcomes across groups.

By Tolulope Fadina, Thorsten Schmidt
arXiv AI
Aug 20

Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Metric for AI Systems

The article argues that for frontier language models, precision—how consistently outputs cluster around the target—should be the key metric rather than capability, which measures average performance. It proposes a simple, non‑circular method to quantify precision by repeatedly scoring deterministic tasks and computing outcome consistency, and demonstrates how this metric can guide decisions about model improvements. The study shows that precision can reveal whether failures are due to systemic misalignment or random noise, and that real‑world measurement is more valuable than rule‑based benchmarks.

By George Andrikopoulos