arXiv:2609.39049v1 Announce Type: cross
Abstract: A large language model (LLM) can rate depression severity directly from a social media post or mark which clinical criteria the post shows and let co...
By Xinkai Chen
A large language model (LLM) can rate depression severity directly from a social media post or mark which clinical criteria the post shows and let code turn the count into a label. The latter is easie...
The authors present the Cross-Platform Fairness Evaluation (CPFE) framework, a five‑axis audit protocol that assesses discriminative performance, calibration, statistical significance, prediction equity, and attribution stability of transformer models. Applying CPFE to four models trained on a Kaggle mental‑health corpus and tested on Reddit and Twitter, they find substantial cross‑platform degradation in AUC (30–40%) and severe calibration failures (ECE rising to 0.5 on Twitter). The study demonstrates that platform‑specific temperature scaling can largely fix calibration without harming discrimination, while prediction equity and attribution stability analyses reveal significant disparities and vocabulary divergence across platforms. The results argue that cross‑platform validation across all CPFE axes should become a standard requirement for mental‑health NLP systems deployed in heterogeneous environments.
By Rajveer Singh Pall, Sameer Yadav
arXiv:2606. 27247v1 Announce Type: new Abstract: In NLP, mental health conditions are often modeled as isolated phenomena, without interpersonal context.
By Parmitha Vangapandu, Sai Ganesh Mokkapati, Sathwik Narkedimilli, MSVPJ Sathvik, Timothy Liu, Simon See, Johannes C. Eichstaedt
The paper introduces TSS (Triple-Stream Stress probe), a diagnostic framework that splits text into lexical, morpho-syntactic, and psycholinguistic style channels to analyze mental health NLP classifiers. Across four English datasets, TSS uncovers a lexical interference effect where adding lexical features harms performance on human-labeled data but not on auto-labeled data, and proposes the Degree of Divergence (DoD) statistic to audit label-source bias. The study demonstrates that style features largely remain effective even after masking content words, emphasizing that shortcut learning is label-source specific rather than clinically relevant.
By Moustafa Yehia Hassan
arXiv:2609.22767v1 Announce Type: cross
Abstract: The IEEE BigData Cup benchmark combines three prediction problems with different output structures: ordinal suicide-risk classification, multi-label...
By Zirui Li, Yanling Li, Kaolanglang Gao