arXiv Machine Learning

Model Multiplicity and Predictive Arbitrariness in Recidivism Risk Assessment

arXiv:2606. 02198v1 Announce Type: new Abstract: Prediction tasks over individual futures, which are inherently noisy, often admit multiple similarly accurate models.

arXiv AI
Sep 2

Measuring consistency via ensemble margin and local prediction variability: Auditing decision systems in the presence of predictive multiplicity

The paper introduces a new consistency criterion for auditing decision systems that combines ensemble margin with local prediction variability to address predictive multiplicity, or the Rashomon effect. It shows that finite ensembles converge to the expected model’s consistency score as ensemble size and sample count grow, and demonstrates that ensembling models from the Rashomon set reduces unchecked incorrect predictions while keeping diversions moderate. Experiments on transformer and fine‑tuned language models for NLP and tabular classification confirm the method’s effectiveness and stronger alignment with existing multiplicity metrics.

By Sinjini Banerjee, Tim Marrinan, Anand D. Sarwate
arXiv Machine Learning
Sep 16

Observational Multiplicity

The paper introduces the concept of observational multiplicity, where multiple probabilistic classifiers can perform similarly yet produce conflicting predictions, undermining interpretability and safety. It proposes measuring this arbitrariness through a regret metric that captures how predictions could shift with different training labels. The authors present a general method to estimate regret, show it varies across dataset groups, and discuss its use for safety via abstention and targeted data collection.

By Erin George, Deanna Needell, Berk Ustun
arXiv AI
Jun 9

Performative Learning Theory

arXiv:2602. 04402v3 Announce Type: replace-cross Abstract: Performative predictions influence the very outcomes they aim to forecast.

By Julian Rodemann, Unai Fischer-Abaigar, James Bailie, Krikamol Muandet
arXiv Machine Learning
Sep 11

A distribution-free certification framework for trustworthy crash-severity prediction

The paper introduces a distribution‑free certification layer that can be applied to any crash‑severity prediction model without modifying the model itself. It provides guarantees for ordinal outcomes, per‑class validity, transfer of coverage to unobserved severities, and one‑sided certificates under deployment shift, all grounded in a functional of the true data law. The framework is evaluated on 5.2 million Texas records, demonstrating a model‑independent lower bound on set width for vulnerable road users and is released as an open‑source package with theorem‑level tests.

By Amir Rafe, Subasish Das
arXiv AI
Sep 23

Toward Auditable and Calibrated AI for Dementia-Related Crash Severity Prediction: A Selective Deferral Framework to Support Human Review

The paper presents a decision‑aware framework for predicting dementia‑related crash severity that emphasizes auditability and selective deferral. Using 4,781 Texas crash records, the authors evaluate several models—including structured, narrative, fusion, calibrated fusion, BERT‑family, and local large‑language‑model baselines—under a stratified 70/15/15 split. The leakage‑controlled Gemma model achieves the highest macro‑F1 of 0.545, while a calibrated fusion model reaches 0.522 macro‑F1 with an expected calibration error of 0.033; selective deferral further improves performance, raising macro‑F1 to 0.573 at 70% coverage and reducing severity cost to 0.577.

By Gaurab Chhetri, Anika Baitullah, Subasish Das