AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation
arXiv:2602. 05088v4 Announce Type: replace Abstract: Millions of people now use generative AI chatbots for psychological support.
arXiv:2602. 05088v4 Announce Type: replace Abstract: Millions of people now use generative AI chatbots for psychological support.
The paper evaluates the safety of conversational AI therapy bots for Generation Alpha, revealing that while these models understand 76‑82% of youth‑specific vocabulary, they correctly assess clinical risk only 64‑72% of the time, creating a significant vocabulary‑comprehension gap. Six failure patterns—such as sarcasm masking, minimization acceptance, and semantic drift—were identified, with compounded errors leading to a 94% miss rate when three or more patterns co‑occur. The authors estimate 146,880 missed crises annually and recommend mandatory human‑in‑the‑loop systems, quarterly youth‑specific validation, transparent performance disclosure, and regulatory oversight for youth‑facing mental health AI.
arXiv:2607. 08625v1 Announce Type: new Abstract: Consumer-facing health chatbots powered by large language models (LLMs) are increasingly used for symptom assessment.
arXiv:2601.09721v2 Announce Type: replace-cross Abstract: Consumer artificial intelligence chatbots are now accessed by hundreds of millions of users seeking health information, yet systematic evalua...
arXiv:2607. 25485v1 Announce Type: new Abstract: Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf.
arXiv:2607. 24817v1 Announce Type: cross Abstract: Digital mental health interventions (DMHIs) offer scalable support, but ensuring they accurately detect users' intent during volatile situations can be challenging.
Anian is a safety‑gated multimodal AI backend designed for perinatal mental‑health support and mindfulness‑intervention routing. It maps user input into a four‑layer hierarchical state representation—emotion, psychosocial constructs, safety risk, and intervention routes—then fuses local and external risk signals to decide whether to generate AI responses or provide fixed safety content. Prototype evaluation on large public corpora showed high classification performance and perfect high‑risk recall in a controlled stress test, though clinical validity remains unestablished.
This scoping review examined 48 studies on generative AI chatbots designed to deliver motivational interviewing (MI). It found that most systems were text‑based and disembodied, with about half incorporating dynamic adaptation, and that safety reporting was inconsistent. While user perceptions were generally positive and many studies reported MI‑consistent interactions, evidence for sustained behavioral or functional change remains limited.
The paper presents a multi‑perspective annotation framework for detecting medical hallucinations in chatbot responses. It combines first‑pass annotators, a large language model acting as a judge (LaJ) for candidate discovery, and two adjudication stages—medical‑expert review and evidence‑based fact‑checking. The study finds that single‑pass labeling undercounts errors, while multi‑pass adjudication improves coverage but still depends on expert judgment and evidence.
The paper investigates how to better detect factual errors, or hallucinations, in long-form medical chatbot responses. It introduces a multi‑perspective annotation workflow that combines first‑pass labeling, a large language model acting as a judge (LaJ) to surface candidate errors, and two adjudication steps—expert medical review and evidence‑based fact‑checking. The study finds that single‑pass benchmarks miss many errors, that LaJ alone is insufficient, and that adjudicators disagree, indicating that multi‑pass adjudication improves coverage but still depends on human judgment and evidence.
arXiv:2512. 01241v4 Announce Type: replace-cross Abstract: Large language models (LLMs) and medical AI tools are routinely used by physicians and patients for medical advice, yet their clinical safety profiles remain poorly characterized.
arXiv:2511.11689v4 Announce Type: replace-cross Abstract: Generative AI chatbots built for mental health could extend access to care, but evidence from real-world use is limited. We report a single-a...