arXiv:2609.39049v1 Announce Type: cross
Abstract: A large language model (LLM) can rate depression severity directly from a social media post or mark which clinical criteria the post shows and let co...
By Xinkai Chen
arXiv:2606. 05970v1 Announce Type: cross Abstract: Large language models are increasingly used for structured extraction from clinical free-text notes, but the sensitivity of their output to upstream configuration choices is less understood than their accuracy on fixed benchmarks.
By Martin Murin
The study compares six approaches—frontier commercial models, fine‑tuned smaller models, and conventional classifiers—for detecting anxiety in Reddit posts. It reveals a significant lexical bias: 69.3% of anxiety‑labelled posts contain the word "anxiety" or a variant, allowing models to perform well via keyword matching rather than true language understanding. After removing these terms, the frontier model still leads (F1 = 0.846), but a 110 M‑parameter domain‑adapted encoder achieves a close score (F1 = 0.831) without external API calls, and lexical dependence varies widely across models.
By Cris Huynh, Arlene Pham
arXiv:2608. 08746v1 Announce Type: new Abstract: Prospective daily symptom tracking is central to premenstrual health assessment, but repeated ordinal forms impose substantial response burden.
By Yifan Wang
The paper introduces TSS (Triple-Stream Stress probe), a diagnostic framework that splits text into lexical, morpho-syntactic, and psycholinguistic style channels to analyze mental health NLP classifiers. Across four English datasets, TSS uncovers a lexical interference effect where adding lexical features harms performance on human-labeled data but not on auto-labeled data, and proposes the Degree of Divergence (DoD) statistic to audit label-source bias. The study demonstrates that style features largely remain effective even after masking content words, emphasizing that shortcut learning is label-source specific rather than clinically relevant.
By Moustafa Yehia Hassan
The study investigates how a large language model, Gemma-3-27B-PT, internally represents depressive symptoms. By applying mechanistic interpretability methods to the model’s residual stream, researchers found that symptom groups are geometrically distinct at layer 21, and that projected symptom vectors align with clinician-annotated rankings across mood, somatic, and suicidality dimensions. Additionally, a single depression vector at this layer can differentiate depressive from non-depressive text with an AUC of 0.789, suggesting a potential emotional valence gate for symptom projection.
By Fangyi Zhu, Ajay Subramanian, Allison Constant, Camille Wang, Ravish Gupta, Corey J. Keller
arXiv:2609.24885v1 Announce Type: new
Abstract: When a language model answers from a curated corpus via graph-based retrieval, a large grounding uplift does not establish reasoning over the retrieved...
By John J. O'Hare
The study examines how large language models (LLMs) predict depression scores from language responses. In a "Mirror" setup, participants answered structured diagnostic interviews that the LLMs used to predict scores, yielding near-perfect predictions. In a "Non-Mirror" setup, participants gave life history interviews; the LLMs still achieved outstanding prediction accuracy, and both conditions correlated similarly with PHQ-9 scores, indicating that the Mirror advantage disappears when predicting an independent measure. Topic modeling showed different depression themes across interview types, suggesting Mirror evaluations are more about reliability than validity and that Non-Mirror approaches may enhance clinical relevance.
By Tong Li, Rasiq Hussain, Mehak Gupta, Joshua R. Oltmanns
The study investigates how large language models (LLMs) assess psychological distress in online posts from six identity‑based communities. Through a perspectivist annotation task, 321 participants provided 9,587 judgments on 1,198 Reddit posts, revealing modest in‑group agreement (OR = 1.18) that varies across communities. When evaluated against these community‑specific labels, open‑weight LLMs consistently over‑estimate distress—achieving only 31–44% accuracy on posts perceived as none‑to‑mild—while newer models like GPT‑5 and Gemini 2.5 Pro show similar inflation, whereas Claude Opus 4 is more conservative.
"whyItMatters":"The findings highlight that miscalibrated distress detection by LLMs can disproportionately impact the very communities they aim to serve, underscoring the need for equitable AI deployment in mental‑health contexts."
By Andrew Aquilina, Xiang Lorraine Li, Yu-Ru Li
arXiv:2609.22767v1 Announce Type: cross
Abstract: The IEEE BigData Cup benchmark combines three prediction problems with different output structures: ordinal suicide-risk classification, multi-label...
By Zirui Li, Yanling Li, Kaolanglang Gao
arXiv:2607.15870v2 Announce Type: replace
Abstract: Human label variation in natural language inference is increasingly treated as signal rather than noise, but how much of it formal semantic structu...
By Haram Choi (University of Bremen)
arXiv:2609.07766v1 Announce Type: cross
Abstract: Assessing suicide risk from social media text is a small-data, high-stakes setting requiring not only severity prediction but also supporting evidenc...
By Shlok Shelat, Shrey Salvi, Souvik Roy, Manas Gaur, Amit Sheth