Evaluation Metrics as Averaged Outcomes of Fair Gambles
arXiv:2401. 14483v4 Announce Type: replace Abstract: In the current practices of machine learning, the evaluation of forecasts has become a cornerstone of scientific progress.
arXiv:2605. 26703v2 Announce Type: replace-cross Abstract: The classic concept of "calibrated forecasts" and its more recent refinement, "calibeating," are defined with respect to the standard quadratic scoring rule.
arXiv:2401. 14483v4 Announce Type: replace Abstract: In the current practices of machine learning, the evaluation of forecasts has become a cornerstone of scientific progress.
The paper introduces rankECE, a new metric for assessing calibration error in predictive models. Unlike the widely used Expected Calibration Error (ECE), rankECE compares predictions with neighboring probability values, offering theoretical guarantees and empirical evidence that it better approximates ECE than traditional binned methods.
The paper investigates the feasibility of exact truthfulness in calibration measures for sequential binary prediction. It proves that exact truthfulness cannot coexist with completeness and soundness, even when outcomes are independent. The authors then provide two reductions that transform any base calibration measure into additively or multiplicatively approximately truthful ones, achieving a multiplicative truthfulness guarantee that improves upon previous results.
arXiv:2607. 12928v1 Announce Type: new Abstract: We study the online binary sequential calibration problem.
The paper argues that traditional weather forecast evaluations, which focus on statistical comparisons between forecasts and observations, do not adequately capture how forecasts influence real-world decisions. It introduces decision calibration, a framework that assesses probabilistic forecast performance from the decision-maker’s perspective. Using this framework, the authors compare a machine learning model to a classical numerical weather prediction model across various weather-dependent decision tasks, finding that forecast-level performance does not reliably predict decision-level outcomes and that model rankings can shift depending on the decision context.
arXiv:2606. 09517v1 Announce Type: new Abstract: As renewable energy integration increases market volatility, probabilistic electricity price forecasting has become essential for effective risk management.
arXiv:2306. 02704v2 Announce Type: replace-cross Abstract: We introduce \emph{Calibrated Stackelberg Games (CSGs)}, a generalization of the standard Stackelberg Games (SGs) framework.
The paper investigates how different proper scoring rules influence the performance and behavior of large language model (LLM) forecasters. Five scoring rules were compared as training objectives for binary forecasts of real-world events, revealing that while they all theoretically incentivize truthful probability reporting, they produce models with varying calibration, probability usage, and bias, information, and noise profiles. The Brier-trained model achieved the lowest Brier score and highest AUC-ROC, whereas the log-trained model achieved the best log score and lowest calibration error, indicating that scoring rule choice can shape both forecast accuracy and error structure.
arXiv:2607. 16229v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as components of agentic systems that observe, plan, and act.
arXiv:2606. 03245v1 Announce Type: cross Abstract: Concepts of calibration formalize the compatibility between probabilistic predictions and the respective outcomes.
arXiv:2606. 10777v1 Announce Type: new Abstract: Uncertainty estimation is critical for deploying machine learning models in high-stakes settings.
arXiv:2606. 18527v1 Announce Type: cross Abstract: U-calibration studies online forecasting algorithms whose predictions can be consumed by any unknown downstream agent, guaranteeing sublinear regret simultaneously for all proper loss functions.