The paper presents a four‑module, data‑driven framework to identify and prioritize robotic process automation (RPA) opportunities in U.S. hospitals. It includes a process taxonomy, an automation suitability index, a tool‑tier selection recommendation, and a return‑on‑investment analysis, all applied to a synthetic portfolio of twenty hospital processes. The authors demonstrate the framework’s robustness through Monte Carlo simulations and discuss governance and future validation steps.
FLARE is a systematic, uncertainty‑aware framework that evaluates the financial and operational implications of adopting AI in healthcare. It integrates fuzzy logic, time‑driven activity‑based costing, and return‑on‑investment analysis to estimate costs of clinical service delivery, AI development and operation, and the economic impact of workflow integration. A case study on AI‑assisted large vessel occlusion detection in the CT stroke pathway demonstrated that FLARE can quantify conventional pathway costs, AI‑related costs, and AI‑enabled savings, identifying a break‑even threshold of about 3,992 patients per year and a positive first‑year ROI at typical stroke volumes of 5,000 patients.
By Jacob Idoko, Siddhartha Paudel, Mariana Bento, Roberto Souza, Gouri Ginde
arXiv:2608. 07627v1 Announce Type: new Abstract: Hospitals are racing to embed AI, while coping with the surge in adaptation of the technology in other industries, into the triage management, documentation, scheduling, and revenue-cycle workflows, yet most deployments remain as fragmented pilots that stall at the edge of production, exposing patients and institutions to operational fragility, ungoverned risk, and mounting technical debt.
By Manideep Dhar, Ritwik Singh, Sharat Chandra Kumar Manikonda
arXiv:2608. 06112v1 Announce Type: new Abstract: Hospitals are rapidly adopting artificial intelligence for triage, imaging, scheduling etc.
By Manideep Dhar, Ritwik Singh, Sharat Chandra Kumar Manikonda
Hospitals are rapidly adopting artificial intelligence for triage, imaging, scheduling etc. , yet most deployments remain isolated point solutions locked inside departmental silos, resulting in duplicated effort, hidden risks, and unrealized enterprise value.
KnowBench is a new benchmark for clinical AI that measures Effort Reduction (ER), the proportion of system-generated clinical work product accepted by clinicians after expert and safety review. The metric is applied uniformly across various administrative tasks—visit notes, billing codes, orders, EHR summarization, patient summaries, and decision support—using the clinician’s review-and-attestation as ground truth. An initial deployment of Knowtex’s models achieved an aggregate ER of 97.99% across more than one million encounters in six months, with specialty-specific ER ranging from 96.8% to 98.9%.
By Jocelyn Kang, Caroline Zhang
arXiv:2609.21841v1 Announce Type: new
Abstract: Frontier language models now produce professional deliverables that expert graders judge to match human work on a substantial share of economically val...
By Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona Zahid
arXiv:2605.16679v3 Announce Type: replace-cross
Abstract: End-to-end automation of realistic healthcare operations stresses three capabilities underrepresented in current benchmarks: policy density,...
By Haolin Chen, Deon Metelski, Leon Qi, Tao Xia, Joonyul Lee, Steve Brown, Kevin Riley, Frank Wang, T. Y. Alvin Liu, Hank Capps MD, Zeyu Tang, Xiangchen Song, Lingjing Kong, Fan Feng, Tianyi Zeng, Zhiwei Liu, Zixian Ma, Hang Jiang, Fangli Geng, Yuan Yuan, Chenyu You, Qingsong Wen, Hua Wei, Yanjie Fu, Yue Zhao, Carl Yang, Biwei Huang, Kun Zhang, Caiming Xiong, Sanmi Koyejo, Eric P. Xing, Philip S. Yu, Weiran Yao
READY or Not: Reliable Enterprise Agent Deployment introduces a framework for qualifying AI agents for enterprise workflows. It measures reliability and operating cost under various oversight policies, selects the minimum‑cost policy that meets a specified reliability target, and statistically qualifies it on held‑out cases. In a clinical audit study, READY revealed that two agents with nearly identical autonomous accuracy required markedly different levels of human review to achieve the same reliability target.
By Veronica Chatrath (Christy), Bryan Zhu (Christy), Jingxuan Fan (Christy), George Pu (Christy), Soham Dinesh Tiwari (Christy), Soham Dan (Christy), Ryan Young (Christy), Yuan (Christy), Li, Yuang Yao, Apaar Shanker, Minglai Yang, Daniel Yue Zhang, Yunzhong He, Ying Liu, Chenguang Wang, Zhijun Yin, Yuan Xue
arXiv:2602. 19502v2 Announce Type: replace Abstract: Agentic AI systems are increasingly capable of autonomous data science workflows, yet clinical prediction tasks demand domain expertise that purely automated approaches struggle to provide.
By Lalitha Pranathi Pulavarthy, Raajitha Muthyala, Aravind V Kuruvikkattil, Zhenan Yin, Rashmita Kudamala, Saptarshi Purkayastha
AI Morbidity and Mortality (AI M&M) is a structured, blameless framework designed to review clinical AI failures. It combines standardized case intake, evidence preservation, investigator reconstruction, tool‑in‑loop attribution, and corrective‑action tracking, classifying each event across four linked dimensions: Trigger, Mechanism, Clinical Pathway, and Corrective Action. The authors demonstrate the framework with five outpatient medication and clinical decision‑support cases, achieving full agreement among reviewers on all classification axes.
By Paulius Mui, Dean F. Sittig, Steve Labkoff, Sanjay Basu
arXiv:2608. 13209v1 Announce Type: cross Abstract: Many operational decisions are sequences of interventions under a cumulative resource limit, such as a maintenance schedule within a crew-hour budget.
By Minkyoung Kim, Beakcheol Jang