arXiv Statistics ML

CHOIR: heterogeneity-aware conformal prediction for crash injury severity across driver safety strata

arXiv Machine Learning
Sep 25

Not All Synthetic Data Are Equal: Expert-Committee Audit Screening for Imbalanced Crash-Injury-Severity Prediction in Automated Driving Systems

The paper introduces Expert-Committee Audit Screening (ECAS), a framework that evaluates the credibility of synthetic minority samples for predicting crash injury severity in automated driving systems. Using real incident data from the NHTSA, ECAS filters generated samples based on label support, boundary separation, committee agreement, and local plausibility, then selects accepted samples via within‑class percentile normalization and Pareto non‑dominated sorting. The best ECAS configuration, combined with normalizing flow augmentation and a TabPFN classifier, outperformed other evidence settings in balanced accuracy, macro‑F1, and minor‑injury recall, and analysis showed ECAS‑accepted samples were better supported by nearby real crashes.

By Zewei Li, Qiaoqiao Ren, Hang Yang, S. C. Wong, Stergios-Aristoteles Mitoulis, Yun Ye
arXiv Machine Learning
Sep 11

A distribution-free certification framework for trustworthy crash-severity prediction

The paper introduces a distribution‑free certification layer that can be applied to any crash‑severity prediction model without modifying the model itself. It provides guarantees for ordinal outcomes, per‑class validity, transfer of coverage to unobserved severities, and one‑sided certificates under deployment shift, all grounded in a functional of the true data law. The framework is evaluated on 5.2 million Texas records, demonstrating a model‑independent lower bound on set width for vulnerable road users and is released as an open‑source package with theorem‑level tests.

By Amir Rafe, Subasish Das
arXiv AI
Sep 23

Toward Auditable and Calibrated AI for Dementia-Related Crash Severity Prediction: A Selective Deferral Framework to Support Human Review

The paper presents a decision‑aware framework for predicting dementia‑related crash severity that emphasizes auditability and selective deferral. Using 4,781 Texas crash records, the authors evaluate several models—including structured, narrative, fusion, calibrated fusion, BERT‑family, and local large‑language‑model baselines—under a stratified 70/15/15 split. The leakage‑controlled Gemma model achieves the highest macro‑F1 of 0.545, while a calibrated fusion model reaches 0.522 macro‑F1 with an expected calibration error of 0.033; selective deferral further improves performance, raising macro‑F1 to 0.573 at 70% coverage and reducing severity cost to 0.577.

By Gaurab Chhetri, Anika Baitullah, Subasish Das
arXiv Machine Learning
Sep 21

Available Guardrails: Certifying Selective Prediction across ML Systems

The paper introduces a method to certify selective prediction in machine learning systems by computing the availability of safety gates through exact-binomial inversion and dynamic programming. It demonstrates that a truth-informed planner can significantly improve mean coverage over naive approaches, and that reallocating error budgets further enhances coverage across diverse applications such as LLM tool‑calling, content moderation, lesion classification, and recommendation. The study highlights the importance of planning and finite‑sample estimation in ensuring reliable, granular deployment of selective predictors.

By Parivesh Priye, Yufeng Wang, Haibin Ling, Michael Chaykowsky
arXiv Machine Learning
Sep 3

PRISM: An Agentic Multi-Model Architecture for Proactive Safety in Autonomous Transportation Systems

PRISM (Proactive Risk Intelligence and Safety Management) is an agentic multi-model architecture designed to shift autonomous transportation safety from reactive crash avoidance to proactive, continuous risk management. It uses inverse crash‑probability modeling to transform binary crash classifiers into dynamic safety scores, and runs three specialized models—trajectory kinematics, environmental risk, and VRU interaction—coordinated by a reinforcement‑learning reasoning layer. Across 1,296 naturalistic driving scenarios, PRISM achieved a mean safety score of 68/100, classified 77.6% of situations as advisory, and flagged 3.8% as near‑misses, with 11% requiring intervention or emergency response, highlighting trajectory risk and VRU proximity as key safety factors.

By Joyjit Roy, Samaresh Kumar Singh, Sushanta Das
arXiv Machine Learning
Sep 24

CS-WCP: Robust Conformal Sets for LLM-Judge Traffic Shifts with Uncertain Group Proportions

CS-WCP introduces confidence‑set weighted conformal prediction to provide robust prediction sets for large‑language‑model judges when deployment traffic shifts the prevalence of task or policy groups. By constructing simultaneous exact intervals for source and target group masses and taking the union over all compatible ratio vectors, CS‑WCP achieves high coverage (mean 0.973) with few failures across 336 constructed traffic shifts, outperforming standard source conformal prediction. The method offers an auditable coverage safeguard under uncertain mixture weights, focusing on conservative tail protection rather than tighter set sizes.

By Ibne Farabi Shihab, Fariya Afrin
arXiv Machine Learning
5d ago

Missingness-Aware Conformal Prediction Under Cross-Hospital Distribution Shift

The paper introduces a missingness‑aware conformal calibration method for mortality prediction that accounts for cross‑hospital distribution shifts. By selecting a measurement on an independent sample, grouping patients by whether that measurement is recorded, and applying Mondrian calibration within each group, the method avoids reusing calibration outcomes. Experiments on eICU and MIMIC‑IV data show that, compared to pooled calibration, it reduces the worst‑group coverage gap by a median of 1.9 percentage points across six settings, though the benefit varies with predictor and hospital.

By Liang You, Dongwen Ou, Hengyu Shi, Siyuan Dai