arXiv Machine Learning

Heckman-Corrected Epistemic Uncertainty: Selection on Unobservables Defeats Importance Weighting

arXiv:2607. 05806v1 Announce Type: new Abstract: Training data for machine learning is routinely collected by a selection process the model never sees: loans are observed only when granted, outcomes only when a test was ordered.

arXiv Machine Learning
Sep 25

Diverse Geometries, Frozen Weights: Robust Heterogeneous Treatment-Effect Estimation via Causal Expert Ensembles

The paper introduces GeoACE, a five‑expert framework for estimating heterogeneous treatment effects that blends a common anchor‑correction estimator with overlap‑aware and outcome‑guided geometries. The ensemble’s task‑level weights are learned from internal validation predictions, frozen before test evaluation, and applied to experts refitted on the full development data. Adding the outcome‑free, overlap‑aware expert O‑Phi‑ACE consistently improves performance across seven benchmarks, achieving the lowest average rank among 11 comparators.

By Ali Haghpanah Jahromi, Mohammad Taheri
arXiv Machine Learning
Sep 24

Learning Risk Scores Robust to Unobserved Confounders

The paper introduces a method for learning risk scores that remain reliable even when historical data contain unobserved confounders. By treating propensity weights as uncertain and applying sensitivity analysis with Wasserstein distributionally robust optimization, the authors formulate a robust learning problem solvable via an exponential cone program. Experiments on semi‑synthetic UCI data show the approach improves calibration by up to 29.2% over traditional benchmarks and 11.1% over the state of the art, without harming other performance metrics.

By Ryan Edmonds, Yingxiao Ye, Sina Aghaei, Andr\'es G\'omez, \c{C}a\u{g}{\i}l Ko\c{c}yi\u{g}it, Phebe Vayanos
arXiv Machine Learning
Jun 18

Not Just How Much, But Where: Decomposing Epistemic Uncertainty into Per-Class Contributions

arXiv:2602. 21160v3 Announce Type: replace-cross Abstract: In safety-critical classification, the cost of failure is often asymmetric, yet Bayesian deep learning summarises epistemic uncertainty with a single scalar, mutual information (MI), that cannot distinguish whether a model's ignorance involves a benign or safety-critical class.

By Mame Diarra Toure, David A. Stephens
arXiv Machine Learning
Aug 19

Detecting and Discriminating Operator Misspecification in Hybrid PDE-Parameter Learning: a Reference-Free Instrument, with Discrimination Bounded In Sample

The paper introduces a reference‑free instrument that, from a single fit and without an oracle, can detect whether a hybrid PDE‑parameter estimator’s assumed operator is misspecified and distinguish this from mere parameter unidentifiability. In a self‑adjoint parabolic inverse problem, the proposed information‑matrix statistic correctly identifies misspecification with low false‑positive rates, while remaining silent when the design is correctly specified but non‑identifiable. The study demonstrates that conventional accuracy checks can miss significant operator errors, and it maps out the instrument’s blind spots and conditions under which its guarantees hold.

By Eric Fock