arXiv:2608.24555v1 Announce Type: cross
Abstract: Prehospital stroke assessment aims to accurately identify stroke symptoms and make rapid decisions through standardized procedures within an extremel...
By Wentao Yang, Zhenye Xu, Ruoyi Li, Musen Zhang, Yao Guo
arXiv:2606. 24960v1 Announce Type: new Abstract: Tailoring stroke rehabilitation requires assessing how movements are organized, not merely if they succeed.
By Tamim Ahmed, Thanassis Rikakis
arXiv:2605.31351v2 Announce Type: replace-cross
Abstract: AI-based Visually Impaired Assistance (VIA) remains challenging, largely due to the high cost of human evaluation. The VLM-as-a-Judge paradig...
By Yi Zhao, Siqi Wang, Zhe Hu, Yushi Li, Jing Li
AI Morbidity and Mortality (AI M&M) is a structured, blameless framework designed to review clinical AI failures. It combines standardized case intake, evidence preservation, investigator reconstruction, tool‑in‑loop attribution, and corrective‑action tracking, classifying each event across four linked dimensions: Trigger, Mechanism, Clinical Pathway, and Corrective Action. The authors demonstrate the framework with five outpatient medication and clinical decision‑support cases, achieving full agreement among reviewers on all classification axes.
By Paulius Mui, Dean F. Sittig, Steve Labkoff, Sanjay Basu
arXiv:2608. 16831v1 Announce Type: new Abstract: Generative pretraining established reusable task representations; later work on language-based task conditioning and in-context learning showed that a fixed model could adapt its behavior from instructions and demonstrations.
By Minh-Ha Nguyen, Cathy Shyr
ObserverBench is a benchmark framework that evaluates whether internal mechanistic estimators—called observers—are suitable for guiding interventions, control, or safety actions in language models. It separates estimation accuracy from the loss incurred by the chosen action, demonstrating that accurate average estimates can still lead to poor decisions. Experiments on GPT‑2‑small, Qwen2.5‑7B, Gemma‑2‑9B‑it, and Qwen3.5‑9B show that observers trained on action loss tend to select lower‑loss actions, while traditional metrics like AUROC can rank monitors differently from deployment loss, highlighting the need for task‑specific evaluation.