A large language model (LLM) can rate depression severity directly from a social media post or mark which clinical criteria the post shows and let code turn the count into a label. The latter is easie...
The study compares six approaches—frontier commercial models, fine‑tuned smaller models, and conventional classifiers—for detecting anxiety in Reddit posts. It reveals a significant lexical bias: 69.3% of anxiety‑labelled posts contain the word "anxiety" or a variant, allowing models to perform well via keyword matching rather than true language understanding. After removing these terms, the frontier model still leads (F1 = 0.846), but a 110 M‑parameter domain‑adapted encoder achieves a close score (F1 = 0.831) without external API calls, and lexical dependence varies widely across models.
By Cris Huynh, Arlene Pham
arXiv:2606. 05970v1 Announce Type: cross Abstract: Large language models are increasingly used for structured extraction from clinical free-text notes, but the sensitivity of their output to upstream configuration choices is less understood than their accuracy on fixed benchmarks.
By Martin Murin
The paper introduces TSS (Triple-Stream Stress probe), a diagnostic framework that splits text into lexical, morpho-syntactic, and psycholinguistic style channels to analyze mental health NLP classifiers. Across four English datasets, TSS uncovers a lexical interference effect where adding lexical features harms performance on human-labeled data but not on auto-labeled data, and proposes the Degree of Divergence (DoD) statistic to audit label-source bias. The study demonstrates that style features largely remain effective even after masking content words, emphasizing that shortcut learning is label-source specific rather than clinically relevant.
By Moustafa Yehia Hassan
arXiv:2609.22767v1 Announce Type: cross
Abstract: The IEEE BigData Cup benchmark combines three prediction problems with different output structures: ordinal suicide-risk classification, multi-label...
By Zirui Li, Yanling Li, Kaolanglang Gao
arXiv:2608. 08746v1 Announce Type: new Abstract: Prospective daily symptom tracking is central to premenstrual health assessment, but repeated ordinal forms impose substantial response burden.
By Yifan Wang
The study investigates how a large language model, Gemma-3-27B-PT, internally represents depressive symptoms. By applying mechanistic interpretability methods to the model’s residual stream, researchers found that symptom groups are geometrically distinct at layer 21, and that projected symptom vectors align with clinician-annotated rankings across mood, somatic, and suicidality dimensions. Additionally, a single depression vector at this layer can differentiate depressive from non-depressive text with an AUC of 0.789, suggesting a potential emotional valence gate for symptom projection.
By Fangyi Zhu, Ajay Subramanian, Allison Constant, Camille Wang, Ravish Gupta, Corey J. Keller
The authors present the Cross-Platform Fairness Evaluation (CPFE) framework, a five‑axis audit protocol that assesses discriminative performance, calibration, statistical significance, prediction equity, and attribution stability of transformer models. Applying CPFE to four models trained on a Kaggle mental‑health corpus and tested on Reddit and Twitter, they find substantial cross‑platform degradation in AUC (30–40%) and severe calibration failures (ECE rising to 0.5 on Twitter). The study demonstrates that platform‑specific temperature scaling can largely fix calibration without harming discrimination, while prediction equity and attribution stability analyses reveal significant disparities and vocabulary divergence across platforms. The results argue that cross‑platform validation across all CPFE axes should become a standard requirement for mental‑health NLP systems deployed in heterogeneous environments.
By Rajveer Singh Pall, Sameer Yadav
The study examines how large language models (LLMs) predict depression scores from language responses. In a "Mirror" setup, participants answered structured diagnostic interviews that the LLMs used to predict scores, yielding near-perfect predictions. In a "Non-Mirror" setup, participants gave life history interviews; the LLMs still achieved outstanding prediction accuracy, and both conditions correlated similarly with PHQ-9 scores, indicating that the Mirror advantage disappears when predicting an independent measure. Topic modeling showed different depression themes across interview types, suggesting Mirror evaluations are more about reliability than validity and that Non-Mirror approaches may enhance clinical relevance.
By Tong Li, Rasiq Hussain, Mehak Gupta, Joshua R. Oltmanns
The study investigates how large language models (LLMs) assess psychological distress in online posts from six identity‑based communities. Through a perspectivist annotation task, 321 participants provided 9,587 judgments on 1,198 Reddit posts, revealing modest in‑group agreement (OR = 1.18) that varies across communities. When evaluated against these community‑specific labels, open‑weight LLMs consistently over‑estimate distress—achieving only 31–44% accuracy on posts perceived as none‑to‑mild—while newer models like GPT‑5 and Gemini 2.5 Pro show similar inflation, whereas Claude Opus 4 is more conservative.
"whyItMatters":"The findings highlight that miscalibrated distress detection by LLMs can disproportionately impact the very communities they aim to serve, underscoring the need for equitable AI deployment in mental‑health contexts."
By Andrew Aquilina, Xiang Lorraine Li, Yu-Ru Li
arXiv:2609.24885v1 Announce Type: new
Abstract: When a language model answers from a curated corpus via graph-based retrieval, a large grounding uplift does not establish reasoning over the retrieved...
By John J. O'Hare
arXiv:2609.07766v1 Announce Type: cross
Abstract: Assessing suicide risk from social media text is a small-data, high-stakes setting requiring not only severity prediction but also supporting evidenc...
By Shlok Shelat, Shrey Salvi, Souvik Roy, Manas Gaur, Amit Sheth