arXiv:2609. 18622v1 Announce Type: new Abstract: Classifiers can make identical predictions yet require labels to compare their selective performance: confidence ranks weight the same errors differently.
By Tetsuji Kuboyama
arXiv:2412. 12807v4 Announce Type: replace-cross Abstract: Selective classification is a powerful tool for automated decision-making in high-risk scenarios, allowing classifiers to act only when confident and abstain when uncertainty is high.
By Mohamed Ndaoud, Peter Radchenko, Bradley Rava
Classifiers can make identical predictions yet require labels to compare their selective performance: confidence ranks weight the same errors differently. We quantify this requirement for the area und...
arXiv:2607. 03999v1 Announce Type: cross Abstract: Estimating heterogeneous treatment effects (CATE) requires simultaneously detecting effect modification and quantifying estimation uncertainty.
By Pantelis Z. Hadjipantelis, Josephine Chiang, Karthik Nagesh
arXiv:2607. 18088v1 Announce Type: new Abstract: Standard evaluation of many recognition systems contains distribution shift by construction, since benchmarks place disjoint conditions in the training and test splits.
By Weijia Han, Lisha Qu
arXiv:2607. 24562v1 Announce Type: new Abstract: Large language models serve heterogeneous populations structured by domain, topic difficulty, and linguistic style.
By Murilo Salem, Lu\'isa B\"ohm, Daniel Pontes, Anderson Ferrugem
arXiv:2404. 02141v5 Announce Type: replace-cross Abstract: In both observational data and randomized control trials, researchers select statistical models to articulate how the outcome of interest varies with combinations of observable covariates.
By Aparajithan Venkateswaran, Anirudh Sankar, Arun G. Chandrasekhar, Tyler H. McCormick
The paper introduces a method to certify selective prediction in machine learning systems by computing the availability of safety gates through exact-binomial inversion and dynamic programming. It demonstrates that a truth-informed planner can significantly improve mean coverage over naive approaches, and that reallocating error budgets further enhances coverage across diverse applications such as LLM tool‑calling, content moderation, lesion classification, and recommendation. The study highlights the importance of planning and finite‑sample estimation in ensuring reliable, granular deployment of selective predictors.
By Parivesh Priye, Yufeng Wang, Haibin Ling, Michael Chaykowsky
The paper derives the exact finite‑population variance of a weighted risk estimator for rare‑event forecasting in dependent sequences and solves for the optimal stratified allocation of a small subsample. It shows that the optimal allocation is equal across strata, independent of the imbalance ratio, and provides a parameter‑free efficiency prediction A(π,f). The authors validate these theoretical predictions on a real‑world dataset of U.S. equities, demonstrating that the predicted ordering of sampling designs matches empirical results.
By Jaskaran Singh
The paper introduces a new estimand for conditional distributional treatment effects that captures how treatments influence the entire outcome distribution, including variance and tail risks, in a covariate-dependent manner. It presents a doubly robust estimator that is minimax optimal locally and uses it to construct a test for global homogeneity of conditional potential outcome distributions. The test accommodates discrepancies beyond the maximum mean discrepancy, guarantees valid type‑1 error, is consistent against fixed alternatives, and includes a computationally efficient, permutation‑free algorithm with exact closed‑form expressions for two natural discrepancies.
By Saksham Jain, Alex Luedtke
arXiv:2609.09245v1 Announce Type: new
Abstract: Repeated-sampling evaluations increasingly extrapolate pass@k far beyond the number n of samples collected per problem. We show that, in the pooled/ran...
By Pranav Singh, Prashant Singh
arXiv:2607. 08347v1 Announce Type: cross Abstract: Active testing provides a label--efficient approach to risk estimation by adaptively selecting which test points should be labelled.
By Kianoosh Ashouritaklimi, Valentin Kilian, Daolang Huang, Tom Rainforth, Fran\c{c}ois Caron