arXiv AI By Gaurab Chhetri, Anika Baitullah, Subasish Das

Toward Auditable and Calibrated AI for Dementia-Related Crash Severity Prediction: A Selective Deferral Framework to Support Human Review

Read the original on arXiv AI →

The paper presents a decision‑aware framework for predicting dementia‑related crash severity that emphasizes auditability and selective deferral. Using 4,781 Texas crash records, the authors evaluate several models—including structured, narrative, fusion, calibrated fusion, BERT‑family, and local large‑language‑model baselines—under a stratified 70/15/15 split. The leakage‑controlled Gemma model achieves the highest macro‑F1 of 0.545, while a calibrated fusion model reaches 0.522 macro‑F1 with an expected calibration error of 0.033; selective deferral further improves performance, raising macro‑F1 to 0.573 at 70% coverage and reducing severity cost to 0.577.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 25

Not All Synthetic Data Are Equal: Expert-Committee Audit Screening for Imbalanced Crash-Injury-Severity Prediction in Automated Driving Systems

The paper introduces Expert-Committee Audit Screening (ECAS), a framework that evaluates the credibility of synthetic minority samples for predicting crash injury severity in automated driving systems. Using real incident data from the NHTSA, ECAS filters generated samples based on label support, boundary separation, committee agreement, and local plausibility, then selects accepted samples via within‑class percentile normalization and Pareto non‑dominated sorting. The best ECAS configuration, combined with normalizing flow augmentation and a TabPFN classifier, outperformed other evidence settings in balanced accuracy, macro‑F1, and minor‑injury recall, and analysis showed ECAS‑accepted samples were better supported by nearby real crashes.

By Zewei Li, Qiaoqiao Ren, Hang Yang, S. C. Wong, Stergios-Aristoteles Mitoulis, Yun Ye
arXiv AI
Jul 10

A safety-oriented hypothetico-deductive framework for AI-assisted differential diagnosis

arXiv:2607. 08038v1 Announce Type: new Abstract: Diagnostic error is a major threat to patient safety, yet current large language model (LLM) systems often treat diagnosis as a one-shot prediction task, lacking safeguards against missed high-risk alternatives or rigorous verification of their reasoning.

By Fan Ma, Mauro Giuffr\`e, Donald Wright, Kent McCann, Mark Iscoe, Lingfei Qian, Mingyang Jiang, Chi Wing Ng, Na Hong, Huan He, Cathy Shyr, Qingyu Chen, Lee Schwamm, Lucila Ohno-Machado, Hua Xu
arXiv Machine Learning
Sep 16

Crash Narrative-Guided Countermeasure Recommendation Using Large Language Models: A Retrieval-Augmented Generation Framework for Intersection Safety

arXiv:2609.15997v1 Announce Type: cross Abstract: Improving safety at intersections requires identifying crash mechanisms and recommending appropriate countermeasures. However, this process tradition...

By Abu Saif Md Nasim Uddin, Mohamed Abdel-Aty, Zubayer Islam, Parvez Anowar, Chenzhu Wang