arXiv Machine Learning

Robust Ambiguity Detection (RAD) From Model- and Feature-Space Consistency

arXiv:2608. 11541v1 Announce Type: new Abstract: Machine learning models should be robust, in the sense of remaining predictively consistent under permissible variations.

Hugging Face Trending Papers
6d ago

Does Model Uncertainty Track Human Ambiguity? Evidence from Multi-Annotator Vision Benchmarks

The paper examines whether model uncertainty aligns with human disagreement on vision tasks. Using multi‑annotator datasets (FER+ and CIFAR‑10H), the authors find that pretrained models rarely reflect the ambiguity humans perceive, with weak correlations between model confidence and human disagreement. Predictive multiplicity offers only modest improvement, indicating that common uncertainty metrics fail to flag ambiguous cases.

arXiv AI
Sep 7

From Answers to Interpretations: Rethinking Ambiguity-Induced Aleatoric Uncertainty Estimation in LLMs

The paper challenges the common practice of estimating aleatoric uncertainty in large language models (LLMs) by generating multiple clarified inputs and comparing the resulting answers. It argues that answers are unnecessary, costly, and can introduce epistemic leakage, proposing instead a clarification-only method that directly assesses ambiguity from the space of plausible interpretations. Experiments on three benchmarks show the new approach improves AUROC, reduces computational cost, and yields uncertainty estimates less correlated with epistemic uncertainty.

By Omer Nahum, Niv Nayman, Jonathan Fhima, Alon Zolfi, Jeremy Levy, Shai Mazor, Paolo Favaro
arXiv Machine Learning
Sep 14

PLSP (Pre-hoc Liminal Space Profiling): OOD Prediction over Detection -- An Anticipatory Approach for Machine Learning Model Reliability

The paper introduces PLSP (Pre-hoc Liminal Space Profiling), an anticipatory framework for predicting out-of-distribution (OOD) data before inference. It proposes a dataset‑independent metric called the CREDibility Score (CREDS) and introduces credibility curves and heat maps to analyze a model’s maximum credibility and behavior across datasets. Experiments on multiple datasets show that CREDS can improve model robustness to OOD prediction.

By Vipul Bansal, Himanshu Buckchash, Balasubramanian Raman, Deepak Dhungana
arXiv Machine Learning
Jul 10

Robustness Quantification for Discriminative Models: a New Robustness Metric and its Application to Dynamic Classifier Selection

arXiv:2603. 23318v2 Announce Type: replace Abstract: Among the different possible strategies for evaluating the reliability of individual predictions of classifiers, robustness quantification stands out as a method that evaluates how much uncertainty a classifier could cope with before changing its prediction.

By Rodrigo F. L. Lassance, Jasper De Bock