arXiv Machine Learning

From Uncertainty to Failure Attribution: Self-Diagnosing Models for Failure Attribution under Distribution Shift

arXiv:2608. 07953v1 Announce Type: new Abstract: Distribution shift poses a significant challenge to the robustness of machine learning models, but the current solutions only aim to detect out-of-distribution (OOD) samples and predict uncertainty levels.

arXiv AI
Aug 28

Diagnosing Conformal Prediction Failures Under Distribution Shift: A COVID-19 Case Study

The paper introduces SHAP concentration as a pre‑deployment diagnostic for detecting when conformal prediction will fail under distribution shift, specifically in gradient‑boosted classifiers. Using a COVID‑19 supply‑chain case study, the authors show that higher feature‑importance concentration correlates with larger drops in coverage, while standard shift detectors cannot differentiate between catastrophic and robust outcomes. The diagnostic is validated on additional datasets, and a formal theorem links concentration to worsening conformity‑score bounds, though it does not capture global‑sensitivity failures in neural networks.

By Chorok Lee
arXiv Machine Learning
Jun 19

Evaluating deep learning models for fault diagnosis of a rotating machinery with epistemic and aleatoric uncertainty

arXiv:2412. 18980v2 Announce Type: replace Abstract: Uncertainty-aware deep learning (DL) models recently gained attention in fault diagnosis as a way to promote the reliable detection of faults when out-of-distribution (OOD) data arise from unseen faults (epistemic uncertainty) or the presence of noise (aleatoric uncertainty).

By Reza Jalayer, Masoud Jalayer, Andrea Mor, Carlotta Orsenigo, Carlo Vercellis
arXiv Machine Learning
Sep 14

PLSP (Pre-hoc Liminal Space Profiling): OOD Prediction over Detection -- An Anticipatory Approach for Machine Learning Model Reliability

The paper introduces PLSP (Pre-hoc Liminal Space Profiling), an anticipatory framework for predicting out-of-distribution (OOD) data before inference. It proposes a dataset‑independent metric called the CREDibility Score (CREDS) and introduces credibility curves and heat maps to analyze a model’s maximum credibility and behavior across datasets. Experiments on multiple datasets show that CREDS can improve model robustness to OOD prediction.

By Vipul Bansal, Himanshu Buckchash, Balasubramanian Raman, Deepak Dhungana