arXiv:2608. 06206v1 Announce Type: cross Abstract: Conformal prediction endows arbitrary black-box predictors with finite-sample, distribution-free marginal coverage, yet marginal validity can hide severe covariate-specific miscalibration, while exact distribution-free conditional coverage is finite-sample unattainable.
By Anton Conrad, Rustam Isaev, Denis Belomestny, Eric Moulines, Sergey Samsonov
arXiv:2606. 13221v2 Announce Type: replace Abstract: Evaluating new large language models typically requires costly human annotation campaigns at scale.
By Bora Kargi, David Salinas
arXiv:2608. 08002v1 Announce Type: new Abstract: Language-model judges and reward models enable scalable supervision, but finite optimization can exploit evaluator errors rather than improve response quality.
By Fariya Afrin, Ibne Farabi Shihab
The paper introduces “ℝD_{CF5}”, a probe‑based estimator that predicts the region‑wise gain of a dynamic ensemble over the best static blend in regression tasks under distribution shift. Across 12 benchmark dataset‑shift pairs, the estimator achieves a Spearman correlation of +0.98 with actual test gains, outperforming alternative diagnostics. The authors also present a Probe‑Validated Ensemble Selector that chooses between a static affine stacker and dynamic realizers, demonstrating risk reductions of up to 16% in prospective deployments.
By Tianxin Zhou, Ruixi Lin
arXiv:2606. 15217v1 Announce Type: cross Abstract: Offline model-based optimization (MBO) proposes candidates by optimizing a surrogate trained on a fixed historical dataset.
By Seungjin Choi
arXiv:2609.17091v1 Announce Type: new
Abstract: Multi-target regression requires a model to simultaneously predict several related outputs. Conformal prediction provides distribution-free, finite-sam...
By Sylvain Rousseau (Heudiasyc), Soundouss Messoudi (Heudiasyc)
arXiv:2608. 02455v1 Announce Type: new Abstract: Human-centered assessment tasks, which are essential for systematic decision-making, rely heavily on human judgment and typically lack verifiable ground truth.
By Zejun Xie, Xintong Li, Guang Wang, Desheng Zhang
The paper argues that traditional global calibration metrics, such as Expected Calibration Error and Brier Score, are confounded by differences in model accuracy when comparing large language models. It introduces ACE, an accuracy‑controlled evaluation framework that offers Instance‑Aligned, Distribution‑Aligned, and Candidate‑Aligned views to provide fairer cross‑model comparisons. Experiments across various benchmarks reveal that many reported calibration advantages disappear after accuracy control and that model rankings often reverse, indicating that raw global metrics are unreliable for cross‑model calibration assessment.
By Zhichao Yang, Caiqi Zhang, Ruihan Yang, Chengzu Li, Nigel Collier, Deqing Yang
arXiv:2608. 03416v1 Announce Type: new Abstract: Large language models (LLMs) are now regularly asked to forecast real-world events, but comparisons are often difficult because models receive different information, use different tools, and are evaluated under different rules.
By Jonaid Shianifar, Iias Faiud
arXiv:2607. 10139v1 Announce Type: cross Abstract: Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each carry a cost: self-consistency inherits the errors of the single model it resamples, and trained reward models need labeled data and transfer poorly off-distribution.
By Ning Liu
arXiv:2608. 16210v1 Announce Type: new Abstract: Aggregate accuracy hides where models succeed and fail.
By Zhi Zhang, Lingfeng Lyu, Yue Kang, Doudou Zhou
arXiv:2606. 09705v1 Announce Type: new Abstract: Scientific generative modeling often requires size transfer, where models trained on small systems are evaluated on larger ones.
By Wenjie Xi