arXiv:2609.22767v1 Announce Type: cross
Abstract: The IEEE BigData Cup benchmark combines three prediction problems with different output structures: ordinal suicide-risk classification, multi-label...
By Zirui Li, Yanling Li, Kaolanglang Gao
arXiv:2609.07766v1 Announce Type: cross
Abstract: Assessing suicide risk from social media text is a small-data, high-stakes setting requiring not only severity prediction but also supporting evidenc...
By Shlok Shelat, Shrey Salvi, Souvik Roy, Manas Gaur, Amit Sheth
arXiv:2609.22696v1 Announce Type: new
Abstract: Decentralized social media platforms create new opportunities and challenges for computational mental health research because data access, moderation,...
By Gaurab Chhetri, Anandi Dutta, Subasish Das
arXiv:2606. 28334v1 Announce Type: cross Abstract: Recent advances in artificial intelligence (AI) and social media data have led to growing optimism about the ability to detect suicide risk at scale.
By Yaakov Ophir, Ofri Hefetz, Refael Tikochinski, Kfir Bar, Shir Lissak, Shulamit Grinapol, Haya Wachtel, Eyal Fruchter, Roi Reichart
The study evaluates large language models for assessing suicide risk in Arabic crisis helpline calls, comparing Arabic and English models. Using de‑identified transcripts from Lebanon’s National Lifeline, the researchers fine‑tuned instruction‑tuned LLMs and transformer encoders, achieving a macro‑F1 of 81.19 and ROC‑AUC of 90.61 for high‑risk calls in Arabic, and 85.00/92.59 in English. The results show that high‑risk calls are more distinguishable than at‑risk calls, and translating to English does not degrade performance, indicating potential for operator‑facing tools.
By Linhai Ma, Rita El Hachem, Mahatab El Hajj, Lilian Ghandour, Samah Fodeh
Humanitarian reports are long, noisy, and multi-topic, making it difficult to consolidate decision-relevant causal evidence. We present a ReliefWeb study (2000-2024) and a two-stage Large Language Model (LLM) pipeline that extracts structured intervention-outcome records with direction and strength attributes.
arXiv:2605.22286v2 Announce Type: replace-cross
Abstract: Text-based counseling provides a valuable source of information for assessing depression severity. We study prediction of the total score on...
By Zhaomin Wu, Jiayi Li, Bingsheng He
arXiv:2512. 06227v3 Announce Type: replace-cross Abstract: Real-world indicators play an important role in many Natural Language Processing (NLP) applications, such as life events for mental health analysis and risky behaviours for online safety, yet labelling such information is often costly and/or difficult due to its multi-label and dynamic nature.
By Junyu Mao, Anthony Hills, Talia Tseriotou, Maria Liakata, Aya Shamir, Dan Sayda, Dana Atzil-Slonim, Natalie Djohari, Pamela Ugwudike, Mahesan Niranjan, Stuart E. Middleton
arXiv:2608. 14649v1 Announce Type: new Abstract: We present dLLM-SetScore, a training-free method that uses discrete masked-diffusion language models for multi-label text classification.
By Pawan Kumar
arXiv:2606. 10380v1 Announce Type: cross Abstract: Real-world crisis intervention is inherently conversational, yet existing research largely focuses on static texts.
By Grace Byun, Abigail Lott, Rebecca Lipschutz, Sean T. Minton, Elizabeth A. Stinson, Jinho D. Choi
arXiv:2605. 23055v2 Announce Type: replace-cross Abstract: Frontier language models sometimes recognize that they are being evaluated and adjust their behavior, undermining validity of benchmark results.
By Changling Li, Terry Jingchen Zhang, Jie Zhang, Zhijing Jin, Sahar Abdelnabi, Maksym Andriushchenko
The authors present the Cross-Platform Fairness Evaluation (CPFE) framework, a five‑axis audit protocol that assesses discriminative performance, calibration, statistical significance, prediction equity, and attribution stability of transformer models. Applying CPFE to four models trained on a Kaggle mental‑health corpus and tested on Reddit and Twitter, they find substantial cross‑platform degradation in AUC (30–40%) and severe calibration failures (ECE rising to 0.5 on Twitter). The study demonstrates that platform‑specific temperature scaling can largely fix calibration without harming discrimination, while prediction equity and attribution stability analyses reveal significant disparities and vocabulary divergence across platforms. The results argue that cross‑platform validation across all CPFE axes should become a standard requirement for mental‑health NLP systems deployed in heterogeneous environments.
By Rajveer Singh Pall, Sameer Yadav