arXiv:2411.08881v3 Announce Type: replace-cross
Abstract: AI-based systems, including Large Language Models (LLMs), impact millions by supporting diverse tasks but face issues like misinformation, bi...
By Jos\'e Antonio Siqueira de Cerqueira, Mamia Agbese, Rebekah Rousi, Nannan Xi, Juho Hamari, Pekka Abrahamsson
AI Morbidity and Mortality (AI M&M) is a structured, blameless framework designed to review clinical AI failures. It combines standardized case intake, evidence preservation, investigator reconstruction, tool‑in‑loop attribution, and corrective‑action tracking, classifying each event across four linked dimensions: Trigger, Mechanism, Clinical Pathway, and Corrective Action. The authors demonstrate the framework with five outpatient medication and clinical decision‑support cases, achieving full agreement among reviewers on all classification axes.
By Paulius Mui, Dean F. Sittig, Steve Labkoff, Sanjay Basu
arXiv:2606. 26298v1 Announce Type: new Abstract: Autonomous AI agents may begin to perform consequential, irreversible actions such as clinical prescribing and production software deployment.
By Jakob Salfeld-Nebgen
arXiv:2606. 14149v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in healthcare settings, yet their tendency to hallucinate poses risks when clinical decisions are involved.
By Muhammad Osama, Maheera Amjad, Zartasha Mustansar, Arslan Shaukat, Muhammad U. S. Khan
arXiv:2607. 17225v1 Announce Type: cross Abstract: Agentic AI systems do not just predict or recommend; they plan, maintain state, and act in external environments with varying degrees of autonomy.
By Chetan Arora, Andreas Vogelsang, Abbi Sharma
Hospitals are rapidly adopting artificial intelligence for triage, imaging, scheduling etc. , yet most deployments remain isolated point solutions locked inside departmental silos, resulting in duplicated effort, hidden risks, and unrealized enterprise value.
The paper proposes AI Deployment Accountability Engineering (ADAE), a new subdiscipline focused on establishing measurable, continuous, and actionable accountability for AI systems once they are deployed. ADAE treats accountability as a deployment-layer property, aiming to ensure systems remain within acceptable risk limits, identify failure contexts, attribute failures across technical and human components, and translate technical failures into downstream consequences. The authors outline a research agenda built around four pillars—structured discovery of context-dependent failure modes, privacy-preserving accountability measurement, system-level risk analysis for agentic AI, and translation of technical failures into operational and institutional risks—to support timely intervention in safety-critical socio-technical environments.
By Murat Kantarcioglu
arXiv:2606. 19899v1 Announce Type: cross Abstract: This paper addresses a rapidly emerging policy challenge: how to generate and interpret credible evidence about the biological capabilities and risks of AI scientists, or agentic AI systems capable of autonomously or collaboratively performing multi-step scientific tasks.
By Patricia Paskov, Jeffrey Lee, Kyle Brady, Alyssa Worland
arXiv:2508. 05132v3 Announce Type: replace-cross Abstract: As medical LLMs transition to clinical deployment, assessing their ethical reasoning capability becomes critical.
By Chang Hong, Minghao Wu, Qingying Xiao, Yuchi Wang, Xiang Wan, Guangjun Yu, Benyou Wang, Yan Hu
arXiv:2606. 30970v1 Announce Type: new Abstract: Autonomous AI agents increasingly perform consequential actions on behalf of human principals, including financial transactions, external communications, and enterprise workflows.
By Anuj Kaul, Qianlong Lan, Pranay Gupta
arXiv:2608. 06112v1 Announce Type: new Abstract: Hospitals are rapidly adopting artificial intelligence for triage, imaging, scheduling etc.
By Manideep Dhar, Ritwik Singh, Sharat Chandra Kumar Manikonda
The paper introduces MIMIC-DOS, a dataset derived from MIMIC-IV that focuses on ICU cases where patient symptoms and medical signs are discordant. It presents CARE, a privacy‑compliant multi‑stage agentic reasoning framework that uses a proprietary LLM to generate structured categories and transitions, while a local LLM performs evidence acquisition and decision‑making. In retrospective evaluations on MIMIC‑DOS, CARE outperforms other LLMs and agentic workflows, demonstrating stronger handling of conflicting clinical evidence while preserving patient privacy.
By Haochen Liu, Weien Li, Rui Song, Zeyu Li, Chun Jason Xue, Xiao-Yang Liu, Sam Nallaperuma-Herzberg, Xue Liu, Ye Yuan