arXiv AI

Beyond Accuracy: Interpreting Topic Representation in Suicide Ideation Detection Models

arXiv:2606. 07714v1 Announce Type: cross Abstract: Suicide ideation detection models are typically evaluated using aggregate performance metrics, yet little is known about how they internally represent psychologically meaningful risk factors.

MIT News AI
Sep 24

Estimating suicide risk from text

A new language‑processing tool has been developed to estimate suicide risk from natural language. By analyzing text, the tool can identify individuals at the highest risk. This capability could enable faster and more targeted interventions.

By Jennifer Michalowski | McGovern Institute for Brain Research
arXiv Machine Learning
1d ago

Explainable Suicide Risk Assessment on Social Media with Multi-Task QLoRA

The paper presents a system for the IEEE BigData 2026 Cup on Explainable Suicide Risk Assessment on Social Media. It tackles three tasks—risk-level classification, evidence phrase extraction, and multi-label factor identification—using Qwen2.5-Instruct models adapted with quantized low-rank adaptation (QLoRA) and an answer-masked causal language-model objective. The final system achieved a composite score of 0.7738, with 0.8089 on Task 1 and 0.6919 on Task 2, demonstrating that task‑specific training and tailored aggregation improve performance across the three tasks.

By Xuan Zhong Feng, Geoffrey Martin, Hexin Dong, Yifan Peng
arXiv AI
Aug 20

Large Language Models in Mental Health: A Systematic Review of Applications, Innovations, and Ethical Challenges

The paper reviews how large language models are applied in mental health, covering areas such as social media analysis, clinical conversational agents, therapy support tools, prompt engineering, and multimodal learning. It synthesizes interdisciplinary studies that use social media posts, electronic medical records, and multimodal inputs to detect depression, assess suicide risk, provide personalized therapy, and generate psychoeducational content. The review also discusses advances in model interpretability, annotation strategies, multimodal fusion techniques, and highlights ethical, sociotechnical, and regulatory challenges while proposing frameworks for safe, equitable, and accountable deployment.

By Yisong Chen, Yifan Gao, Sijing Yu, Chuqing Zhao, Yang Lu
arXiv Machine Learning
Aug 12

Automated Data Enrichment using Confidence-Aware Fine-Grained Debate among Open-Source LLMs for Mental Health and Online Safety

arXiv:2512. 06227v3 Announce Type: replace-cross Abstract: Real-world indicators play an important role in many Natural Language Processing (NLP) applications, such as life events for mental health analysis and risky behaviours for online safety, yet labelling such information is often costly and/or difficult due to its multi-label and dynamic nature.

By Junyu Mao, Anthony Hills, Talia Tseriotou, Maria Liakata, Aya Shamir, Dan Sayda, Dana Atzil-Slonim, Natalie Djohari, Pamela Ugwudike, Mahesan Niranjan, Stuart E. Middleton
arXiv AI
Sep 1

Label Semantic Expansion via Label Guided Neural Topic Modeling

The paper introduces Label Semantic Expansion (LSE), a method that enriches sparse label representations by adding descriptive topic words grounded in a corpus. It proposes a Label-Guided Neural Topic Model (LGNTM) that learns label-aligned topics, integrates lexical and document semantics, and maintains consistency between topic and label structures. Experiments show that LSE and LGNTM improve label-topic alignment, label expansion, topic quality, and downstream classification performance.

By Haojia Zheng, Yuyin Lu, Juntian Huang, Fan Ou, Yanghui Rao, Haoran Xie, Fu Lee Wang
arXiv Computation and Language
Sep 10

Beyond Top Words: MonoTM for Topic Modeling with Interpretable Monosemantic Features

MonoTM is an interpretable topic modeling framework that separates the estimation of document–topic mixtures from the generation of topic descriptors. It uses sparse autoencoders to extract dense, interpretable features for mixture estimation, then learns topic descriptors from a distinct set of corpus‑grounded semantic features. This approach preserves global topic structure while providing more meaningful, semantic‑unit descriptors than traditional top‑word lists.

By Una Joh, Bei Yu