arXiv Machine Learning

Who Wins Where? Conformal Model Comparison for Local Superiority

arXiv:2607. 29053v1 Announce Type: new Abstract: Standard model comparison is global, aggregating losses across the covariate space to declare a single winner.

arXiv Machine Learning
Aug 7

Beyond Marginal Validity: Finite-Sample Guarantees for Localized Conformal Prediction

arXiv:2608. 06206v1 Announce Type: cross Abstract: Conformal prediction endows arbitrary black-box predictors with finite-sample, distribution-free marginal coverage, yet marginal validity can hide severe covariate-specific miscalibration, while exact distribution-free conditional coverage is finite-sample unattainable.

By Anton Conrad, Rustam Isaev, Denis Belomestny, Eric Moulines, Sergey Samsonov
arXiv Machine Learning
Aug 20

When Does Dynamic Ensembling Pay Off? Diagnosing Regionwise Gains in Regression under Distribution Shift

The paper introduces “ℝD_{CF5}”, a probe‑based estimator that predicts the region‑wise gain of a dynamic ensemble over the best static blend in regression tasks under distribution shift. Across 12 benchmark dataset‑shift pairs, the estimator achieves a Spearman correlation of +0.98 with actual test gains, outperforming alternative diagnostics. The authors also present a Probe‑Validated Ensemble Selector that chooses between a static affine stacker and dynamic realizers, demonstrating risk reductions of up to 16% in prospective deployments.

By Tianxin Zhou, Ruixi Lin
arXiv Computation and Language
Sep 1

When Calibration Rankings Reverse: Accuracy-Controlled Evaluation for Fair Comparison of LLMs

The paper argues that traditional global calibration metrics, such as Expected Calibration Error and Brier Score, are confounded by differences in model accuracy when comparing large language models. It introduces ACE, an accuracy‑controlled evaluation framework that offers Instance‑Aligned, Distribution‑Aligned, and Candidate‑Aligned views to provide fairer cross‑model comparisons. Experiments across various benchmarks reveal that many reported calibration advantages disappear after accuracy control and that model rankings often reverse, indicating that raw global metrics are unreliable for cross‑model calibration assessment.

By Zhichao Yang, Caiqi Zhang, Ruihan Yang, Chengzu Li, Nigel Collier, Deqing Yang
arXiv AI
Jul 14

LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning

arXiv:2607. 10139v1 Announce Type: cross Abstract: Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each carry a cost: self-consistency inherits the errors of the single model it resamples, and trained reward models need labeled data and transfer poorly off-distribution.

By Ning Liu