arXiv:2609. 18622v1 Announce Type: new Abstract: Classifiers can make identical predictions yet require labels to compare their selective performance: confidence ranks weight the same errors differently.
By Tetsuji Kuboyama
arXiv:2609.37493v1 Announce Type: cross
Abstract: Serving an answer from a large language model requires deciding when to abstain, yet a verifier's ranking accuracy alone does not determine the error...
By Dongyub Jude Lee, Jungseob Lee, Chanjun Park, Hyeonseok Moon, Heuiseok Lim
arXiv:2607. 18088v1 Announce Type: new Abstract: Standard evaluation of many recognition systems contains distribution shift by construction, since benchmarks place disjoint conditions in the training and test splits.
By Weijia Han, Lisha Qu
arXiv:2609.06873v1 Announce Type: cross
Abstract: We study how a limited labeling budget should be allocated to minimize multiclass zero-one classification risk. We consider parametric classification...
By F. Setoudehtanzangi, Geoffrey J. McLachlan
arXiv:2606. 15153v1 Announce Type: new Abstract: Selective prediction with distribution-free risk control promises that, with confidence 1-delta over the calibration draw, the error rate of accepted inputs stays below a user budget alpha.
By Jingwen Zhou, Mingzhe Wang
StepCOPS is a new method for selecting a language‑model policy from many checkpoints, prompts, and decoding rules by providing closed‑testing lower‑tail certificates. It uses an independent proposal split to nominate a lower‑tail floor for each candidate, applies exact binomial tests on a fresh certification split, and employs Holm’s step‑down procedure to certify a set of floors. In experiments across 24 configurations and 11 benchmarks, StepCOPS achieves 96.4% selected‑policy coverage, raises the certified floor by 1.5 points over prior methods, stays 0.6 points below a large‑reference jury oracle, and abstains in 2.4% of trials.
By Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma
arXiv:2608. 15565v1 Announce Type: new Abstract: Experience-learning agents for optimization modeling improve by storing verified skills, but existing learners admit knowledge by checking against known answers, which real ticket streams do not provide.
By Junbo Jacob Lian, Huiling Chen, Hanzhang Qin, Chung-Piaw Teo
arXiv:2607. 14157v1 Announce Type: cross Abstract: Retrieval over corpora that mix several domains often returns relevant but wrong-domain evidence that ranking metrics miss and that conformal risk control bounds only marginally, under-covering the worst domains.
By Jayakumar Manoharan
arXiv:2608. 19376v1 Announce Type: cross Abstract: Split-conformal prediction provides marginal coverage under exchangeability and is increasingly used as an abstention layer for zero-shot vision-language models (VLMs).
By Jai Kumar Sharma, Amartya Dutta
The paper introduces a label‑free method called AURCC for selecting the best foundational model for medical image classification when the target domain lacks labels. AURCC uses a pseudo‑label discrepancy computed by the SUDO framework to score models without fine‑tuning. Experiments on chest X‑ray data across three inter‑hospital shifts show that AURCC closely matches the true model ranking, outperforming simple source‑accuracy baselines especially when source data are limited.
By Juan I\~naki Larrea, Lucas Mansilla, Enzo Ferrante
arXiv:2609.17545v1 Announce Type: new
Abstract: Deep learning models for cervical cytology are almost always evaluated as if every prediction must be acted upon, yet a screening system deployed along...
By Nisreen Albzour, Sarah S. Lam
arXiv:2606. 15910v2 Announce Type: replace Abstract: A vision-language model can answer a question about a chest radiograph or a pathology slide fluently and confidently while barely using the image, relying instead on language priors.
By Reza Khanmohammadi, Kundan Thind, Mohammad M. Ghassemi