arXiv Machine Learning

Proper Calibeating

arXiv:2605. 26703v2 Announce Type: replace-cross Abstract: The classic concept of "calibrated forecasts" and its more recent refinement, "calibeating," are defined with respect to the standard quadratic scoring rule.

arXiv Machine Learning
Sep 14

A Ranking Approach for Measuring Calibration

The paper introduces rankECE, a new metric for assessing calibration error in predictive models. Unlike the widely used Expected Calibration Error (ECE), rankECE compares predictions with neighboring probability values, offering theoretical guarantees and empirical evidence that it better approximates ECE than traditional binned methods.

By Anirban Chatterjee, Rina Foygel Barber
arXiv Machine Learning
Aug 24

Truthful Calibration Measures for Sequential Prediction

The paper investigates the feasibility of exact truthfulness in calibration measures for sequential binary prediction. It proves that exact truthfulness cannot coexist with completeness and soundness, even when outcomes are independent. The authors then provide two reductions that transform any base calibration measure into additively or multiplicatively approximately truthful ones, achieving a multiplicative truthfulness guarantee that improves upon previous results.

By Anagha Gokul, Jason Hartline, Lunjia Hu, Jonathan Ullman, Yifan Wu
arXiv Machine Learning
Sep 7

Forecast Skill Is Not Decision Skill: Evidence from Weather-Dependent Decision Tasks

The paper argues that traditional weather forecast evaluations, which focus on statistical comparisons between forecasts and observations, do not adequately capture how forecasts influence real-world decisions. It introduces decision calibration, a framework that assesses probabilistic forecast performance from the decision-maker’s perspective. Using this framework, the authors compare a machine learning model to a classical numerical weather prediction model across various weather-dependent decision tasks, finding that forecast-level performance does not reliably predict decision-level outcomes and that model rankings can shift depending on the decision context.

By Kornelius Raeth, Nicole Ludwig
arXiv Machine Learning
Aug 31

How Proper Scoring Rules Shape LLM Forecasting

The paper investigates how different proper scoring rules influence the performance and behavior of large language model (LLM) forecasters. Five scoring rules were compared as training objectives for binary forecasts of real-world events, revealing that while they all theoretically incentivize truthful probability reporting, they produce models with varying calibration, probability usage, and bias, information, and noise profiles. The Brier-trained model achieved the lowest Brier score and highest AUC-ROC, whereas the log-trained model achieved the best log score and lowest calibration error, indicating that scoring rule choice can shape both forecast accuracy and error structure.

By Benjamin Turtel, Paul Wilczewski, Kris Skotheim, Ville A. Satop\"a\"a, Philip E. Tetlock
arXiv Machine Learning
Jun 18

Toward Simultaneously Optimal Regret in U-Calibration

arXiv:2606. 18527v1 Announce Type: cross Abstract: U-calibration studies online forecasting algorithms whose predictions can be consumed by any unknown downstream agent, guaranteeing sublinear regret simultaneously for all proper loss functions.

By Rafael Frongillo, Haipeng Luo, Nishant A. Mehta, Jon Schneider