arXiv:2607. 21806v1 Announce Type: new Abstract: Predictive machine learning (ML) models are increasingly used to aid human decision-makers across various high-risk domains such as healthcare and criminal justice.
By Jonathan Zhang, Erik Skalnes, Jacob Chen, Michael Oberst
The paper introduces a new consistency criterion for auditing decision systems that combines ensemble margin with local prediction variability to address predictive multiplicity, or the Rashomon effect. It shows that finite ensembles converge to the expected model’s consistency score as ensemble size and sample count grow, and demonstrates that ensembling models from the Rashomon set reduces unchecked incorrect predictions while keeping diversions moderate. Experiments on transformer and fine‑tuned language models for NLP and tabular classification confirm the method’s effectiveness and stronger alignment with existing multiplicity metrics.
By Sinjini Banerjee, Tim Marrinan, Anand D. Sarwate
arXiv:2505. 08908v3 Announce Type: replace-cross Abstract: Many researchers apply classical statistical decision theory to evaluate treatment choices and learn optimal policies.
By Benedikt Koch, Kosuke Imai
arXiv:2608. 12477v1 Announce Type: new Abstract: Clinical prediction models are often developed as if the outcome of interest were cleanly observed for every patient.
By Xiaobin Shen, Chloe Y. H. Huang, Jonathan Elmer, George H. Chen
The paper introduces the concept of observational multiplicity, where multiple probabilistic classifiers can perform similarly yet produce conflicting predictions, undermining interpretability and safety. It proposes measuring this arbitrariness through a regret metric that captures how predictions could shift with different training labels. The authors present a general method to estimate regret, show it varies across dataset groups, and discuss its use for safety via abstention and targeted data collection.
By Erin George, Deanna Needell, Berk Ustun
arXiv:2606. 02671v1 Announce Type: cross Abstract: Machine learning predictors have become essential tools for guiding automated decision making.
By Itai Zilberstein, Ioannis Anagnostides, Tuomas Sandholm