arXiv:2607. 07229v1 Announce Type: new Abstract: Prior work has shown that chain-of-thought (CoT) reasoning is often unfaithful: a model's stated reasoning does not reliably reflect the process that produced its output.
By Silvia Santano
arXiv:2607. 22692v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for emotional support despite lacking mechanisms to safely govern evolving mental health risk.
By Anabela C. Areias, Catarina Botelho, Ant\'onio Farinhas, Areti Vassilopoulos, Dora Janela, Xin Tong, Nuno M. Guerreiro, Maya D'Eon, Fab\'iola Costa, Ricardo Rei
AI Morbidity and Mortality (AI M&M) is a structured, blameless framework designed to review clinical AI failures. It combines standardized case intake, evidence preservation, investigator reconstruction, tool‑in‑loop attribution, and corrective‑action tracking, classifying each event across four linked dimensions: Trigger, Mechanism, Clinical Pathway, and Corrective Action. The authors demonstrate the framework with five outpatient medication and clinical decision‑support cases, achieving full agreement among reviewers on all classification axes.
By Paulius Mui, Dean F. Sittig, Steve Labkoff, Sanjay Basu
arXiv:2607. 01153v1 Announce Type: cross Abstract: Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model has followed an instruction, refused appropriately, complied with a policy, resisted an embedded command, or misreported progress in an agentic task.
By Brett Reynolds
arXiv:2605. 28591v2 Announce Type: replace-cross Abstract: The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings.
By Katharina Deckenbach, Haritz Puerto, Jonas Geiping, Sahar Abdelnabi
Anian is a safety‑gated multimodal AI backend designed for perinatal mental‑health support and mindfulness‑intervention routing. It maps user input into a four‑layer hierarchical state representation—emotion, psychosocial constructs, safety risk, and intervention routes—then fuses local and external risk signals to decide whether to generate AI responses or provide fixed safety content. Prototype evaluation on large public corpora showed high classification performance and perfect high‑risk recall in a controlled stress test, though clinical validity remains unestablished.
By Lei Wang, Xiao Wang, Lei Li
arXiv:2607. 13036v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for decision support in healthcare, but clinical evidence is often incomplete or evolving.
By Oriana Presacan, Andreea Grama, Larisa Irimin\u{a}, Alireza Nik, Jaya Ojha, Vajira Thambawita, Ciprian I. B\u{a}cil\u{a}, Bogdan Ionescu, Michael A. Riegler
arXiv:2602. 05088v4 Announce Type: replace Abstract: Millions of people now use generative AI chatbots for psychological support.
By Kate H. Bentley, Luca Belli, Adam M. Chekroud, Emily J. Ward, Emily R. Dworkin, Emily Van Ark, Kelly M. Johnston, Will Alexander, Millard Brown, Matt Hawrilenko
The paper introduces a clinician‑grounded evaluation platform called InterviewPlayground, which uses a memory‑augmented patient simulator to assess AI‑assisted psychiatric intake systems. It supports comparison across different interviewing styles, reduces clinician workload, and measures clinically relevant performance. In a pilot study, a GPT‑based intake interviewer captured more relevant items but made more unfounded inferences and missed safety concerns compared to clinicians.
By King Shi, Amanda Li, Jonathan Ivey, Synthia Qia Wang, Guan Gui, Hyunseo Kim, Peter Zandi, Jason Straub, Jacob Taylor, Ananya Joshi
arXiv:2606. 05463v1 Announce Type: new Abstract: Patient safety event triage, determining whether a clinical event is reportable under jurisdiction-specific policy, is a high-stakes task typically performed manually by patient safety experts.
By Keqi Han, Ryan Young, Annabel Strauss, Lindsey Hughes, Katharine M. Nesbitt, Nicole Schueler, Che Ngufor, Carl Yang, Yuan Xue, Zhijun Yin
Large language models (LLMs) have become significant providers of mental health support, yet they remain products of an attention economy whose operational and commercial targets favour sustained engagement over the friction that effective psychological support often requires. Developers' safety responses have been largely reactive, addressing the most visible and acute harms while subtler, longer-term patterns of risk (e.
arXiv:2605. 02050v2 Announce Type: replace-cross Abstract: This work establishes a framework for standardizing AI evaluation RCTs (sometimes called human uplift studies).
By Christopher Kelly, Angelica Chowdhury, Alexandra Campili, Bimpe Ayoola, Devin Barbour, Thomas Chen Dawson, Ze Shen Chin, Rokas Gipi\v{s}kis