arXiv Computer Vision

Reliability-Aware Checkpoint Selection for Domain Generalization

arXiv Machine Learning
Jun 19

Model soups need only one ingredient

arXiv:2602. 09689v2 Announce Type: replace Abstract: Fine-tuning large pre-trained models on a target distribution often improves in-distribution (ID) accuracy, but at the cost of out-of-distribution (OOD) robustness as representations specialize to the fine-tuning data.

By Alireza Abdollahpoorrostam, Nikolaos Dimitriadis, Adam Hazimeh, Pascal Frossard
arXiv AI
3d ago

Evaluating Persistent Calibration under Evolving Model Knowledge

The paper introduces the concept of persistent calibration, which requires a confidence estimator to accurately reflect a model’s evolving knowledge without additional supervision. It evaluates this by comparing confidence estimators trained on earlier checkpoints to their performance on later checkpoints using knowledge contrast sets—questions that shift from correct to incorrect answers across checkpoints. The study finds that standard inference-time and fine-tuning methods underperform compared to oracle methods, and suggests that multi-checkpoint training can improve calibration by identifying robust confidence features.

By Victor Wang, Thomas Hofweber, Mohit Bansal, Elias Stengel-Eskin
arXiv Machine Learning
Sep 21

Available Guardrails: Certifying Selective Prediction across ML Systems

The paper introduces a method to certify selective prediction in machine learning systems by computing the availability of safety gates through exact-binomial inversion and dynamic programming. It demonstrates that a truth-informed planner can significantly improve mean coverage over naive approaches, and that reallocating error budgets further enhances coverage across diverse applications such as LLM tool‑calling, content moderation, lesion classification, and recommendation. The study highlights the importance of planning and finite‑sample estimation in ensuring reliable, granular deployment of selective predictors.

By Parivesh Priye, Yufeng Wang, Haibin Ling, Michael Chaykowsky