Evaluation Metrics as Averaged Outcomes of Fair Gambles
arXiv:2401. 14483v4 Announce Type: replace Abstract: In the current practices of machine learning, the evaluation of forecasts has become a cornerstone of scientific progress.
Goodhart’s law famously says: “When a measure becomes a target, it ceases to be a good measure. ” Although originally from economics, it’s something we have to grapple with at OpenAI when figuring out how to optimize objectives that are difficult or costly to measure.
arXiv:2401. 14483v4 Announce Type: replace Abstract: In the current practices of machine learning, the evaluation of forecasts has become a cornerstone of scientific progress.
Naoki Egami has become a standout in political methodology, helping refine tools that give scholars durable results.
arXiv:2607. 17679v1 Announce Type: cross Abstract: This is an invited commentary on the Psychometrika focus article "Fairness Issues and Evaluation in Psychometrics and AI/ML: What Can We Learn from Each Field?
arXiv:2601. 10600v2 Announce Type: replace-cross Abstract: In the context of multi-agent multi-armed bandits (MA-MAB), fairness is often reduced to outcomes: maximizing welfare, reducing inequality, or balancing utilities.
The article discusses how artificial intelligence is reshaping measurement in economics by converting unstructured data into structured variables at low cost, enabling large‑scale measurement that was previously infeasible. It outlines three stages—discovery, construct definition, and observation—where AI impacts the measurement pipeline and stresses the importance of rigorous validation to ensure credible inference. The review offers guidance on navigating the shift from a single scalable measure to multiple plausible ones that can lead to differing empirical conclusions.
arXiv:2605. 26703v2 Announce Type: replace-cross Abstract: The classic concept of "calibrated forecasts" and its more recent refinement, "calibeating," are defined with respect to the standard quadratic scoring rule.
arXiv:2503. 07811v3 Announce Type: replace-cross Abstract: The theory of optimal transportation has developed into a powerful and elegant framework for comparing probability distributions, with wide-ranging applications in all areas of science.
arXiv:2604.25580v2 Announce Type: replace Abstract: Perspective API closes at the end of 2026, removing the de facto standard for toxicity measurement and exposing researchers' dependence on a tool t...
The paper argues that traditional probabilistic fairness metrics can miss significant disparities in the actual consequences of decisions. By introducing a utility-based framework, the authors show that a process can satisfy ε-fairness yet still be maximally unfair when utilities are considered. They apply this framework to college admissions and credit‑risk assessment, demonstrating that equalizing probabilities alone may mask unequal utility outcomes across groups.
arXiv:2610.07374v1 Announce Type: new Abstract: Omniprediction is a learning guarantee which requires a single predictor to be competitive relative to the best hypothesis from a benchmark class for a...
arXiv:2607. 00913v1 Announce Type: new Abstract: As exponential compute scaling continues, will the capabilities of frontier AI models outstrip what is accessible to developers on a small fixed budget?
The article argues that for frontier language models, precision—how consistently outputs cluster around the target—should be the key metric rather than capability, which measures average performance. It proposes a simple, non‑circular method to quantify precision by repeatedly scoring deterministic tasks and computing outcome consistency, and demonstrates how this metric can guide decisions about model improvements. The study shows that precision can reveal whether failures are due to systemic misalignment or random noise, and that real‑world measurement is more valuable than rule‑based benchmarks.