arXiv AI By Rohith Reddy Bellibatlu, Manpreet Singh, Yash Jajoo, Shyamal Lakhanpal, Abhishek Israni

RISED: A Pre-Deployment Evaluation Framework for High-Stakes AI Decision-Support Systems, with Application to Healthcare

Read the original on arXiv AI →

arXiv:2605. 12895v2 Announce Type: replace-cross Abstract: Clinical decision-support systems are expert systems whose recommendations clinicians act on directly, yet they are usually cleared on one aggregate accuracy number from a held-out test set.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 3

The Ceiling Is in the Channel: Auditing Learner Gaps and Measurement Frontiers in Clinical Prediction

The paper introduces a framework that distinguishes two causes of saturation in clinical prediction: a learner gap, where the model fails to use available information, and a measurement‑channel ceiling, where the recorded variables limit performance. It provides theoretical characterizations, finite‑sample diagnostics, and empirical audits across three large cohorts, showing that well‑tuned models approach the frontier while deficient learners leave large gaps. A PRISMA‑guided synthesis across 104 tasks reveals consistent channel‑level patterns, suggesting that improving the learner or the measurement channel can audit and potentially lift performance.

By Sayeed Shafayet Chowdhury, Nusrat Jahan, Snehasis Mukhopadhyay, Shiaofen Fang, Vijay R. Ramakrishnan
arXiv AI
Sep 11

Safe to Stop? Risk-Constrained Stopping for Sequential Clinical Diagnosis Agents

The paper introduces Cros, a risk‑constrained stopping layer for sequential clinical diagnosis agents that determines when to stop testing and make a diagnosis. Cros combines state‑wise error ranking, policy design on disjoint development splits, and exact tests of selective diagnostic error to provide finite‑sample guarantees. On a MIMIC‑derived abdominal‑pain benchmark, Cros achieves higher state‑error AUROC and lower selective error rates compared to baseline stopping methods, though its performance varies across development resplits.

By Yuexin Wu, Vasile Rus