arXiv AI

Before the Labels: How Dataset Construction Shapes Suicidality Detection in Clinical Text

arXiv:2606. 19637v1 Announce Type: cross Abstract: Clinical NLP increasingly relies on electronic health record (EHR) data to detect suicidal behaviors, treating clinical documentation as more reliable ground truth than social media.

arXiv AI
Aug 20

Large Language Models in Mental Health: A Systematic Review of Applications, Innovations, and Ethical Challenges

The paper reviews how large language models are applied in mental health, covering areas such as social media analysis, clinical conversational agents, therapy support tools, prompt engineering, and multimodal learning. It synthesizes interdisciplinary studies that use social media posts, electronic medical records, and multimodal inputs to detect depression, assess suicide risk, provide personalized therapy, and generate psychoeducational content. The review also discusses advances in model interpretability, annotation strategies, multimodal fusion techniques, and highlights ethical, sociotechnical, and regulatory challenges while proposing frameworks for safe, equitable, and accountable deployment.

By Yisong Chen, Yifan Gao, Sijing Yu, Chuqing Zhao, Yang Lu
arXiv Machine Learning
Aug 12

Automated Data Enrichment using Confidence-Aware Fine-Grained Debate among Open-Source LLMs for Mental Health and Online Safety

arXiv:2512. 06227v3 Announce Type: replace-cross Abstract: Real-world indicators play an important role in many Natural Language Processing (NLP) applications, such as life events for mental health analysis and risky behaviours for online safety, yet labelling such information is often costly and/or difficult due to its multi-label and dynamic nature.

By Junyu Mao, Anthony Hills, Talia Tseriotou, Maria Liakata, Aya Shamir, Dan Sayda, Dana Atzil-Slonim, Natalie Djohari, Pamela Ugwudike, Mahesan Niranjan, Stuart E. Middleton
arXiv Computation and Language
Aug 25

Clinically Grounded Privacy Evaluation of Medical LMs

The paper introduces a clinically grounded privacy evaluation framework for medical language models, assessing leakage across a spectrum of adversarial access levels—from publicly inferable demographics to leaked note fragments. Using this framework on an LM pretrained on 378,000 clinical notes, the authors find that routine encounter metadata leads to high verbatim memorization and significant recovery of sensitive diagnoses (e.g., AUROC 0.91 for abortion, 0.82 for HIV). They also note that exact-match memorization can overstate disclosure, with 36% of memorized tokens being templated documentation, underscoring the risks of training on longitudinal clinical data and offering a reusable evaluation tool.

By Sasha Ronaghi, Sana Tonekaboni, Lena Stempfle, Vivian Utti, Jordan Li Cahoon, Nathaniel Hendrix, Ayin Vala, Marzyeh Ghassemi, Emily Alsentzer
arXiv AI
Sep 18

CliniCIRCA: A Modular LLM Framework for Constructing Longitudinal Mental Health Patient Journeys from Raw EHR Narratives

CliniCIRCA is a modular large‑language‑model framework that reconstructs longitudinal mental‑health patient journeys from raw electronic health record narratives. It temporally classifies clinical events in unstructured discharge summaries without explicit timestamps, producing 15,891 tagged events from 52 summaries and correcting 629 errors to create verified gold‑standard timelines. The framework then generates temporally grounded summaries, compressing each source by 1.52×, and scales to produce 1,000 silver‑standard timelines for training, showing that instruction tuning improves event extraction, temporal tagging, and summarization across models.

By Aiwei Ivy Zhang, Nimra Ishfaq, Mohit Chandra, Santiago Alvarez Lesmes, Adam Coscia, Khatiya Chelidze Moon, Xiaohan Ding, Munmun De Choudhury
arXiv AI
Jun 26

Knowledge-augmented Agentic AI for Mental Health Medication Information Seeking

arXiv:2606. 26205v1 Announce Type: new Abstract: Patients increasingly seek medication information online, yet safety knowledge for psychiatric drugs is split between regulatory adverse-event records, which are authoritative but abstract, and patient narratives, which are experience-near but unvalidated.

By Huizi Yu, Jian Liu, Wenkong Wang, Lingyao Li, Jiayan Zhou, Zhaoqian Xue, Xiang Li, Xinxin Lin, Zhiying Liang, Zhuoru Wu, Siyuan Ma, Xin Ma, Lizhou Fan
arXiv Computation and Language
Sep 25

Clinical Intent Extraction: A FHIR-Aligned Representation and the CIRCA Benchmark

The paper introduces Clinical Intent Extraction (CIE), a task that transforms fragmented clinical action annotations into complete structured records called Clinical Intent Representation (CIR). CIR decomposes each action into verb, type, coded target, timing, condition, request‑intent (aligned to HL7 FHIR) and modality, adding dimensions absent in prior datasets. By re‑expressing five heterogeneous corpora into CIR, the authors create CIRCA, a benchmark of 10,011 harmonized intents with human‑validated subsets, crosswalks, and a deterministic FHIR R4 mapper, and demonstrate that existing models perform poorly on the full task, highlighting the need for targeted development.

By Alexander Apartsin, Yehudit Aperstein