arXiv:2609.39049v1 Announce Type: cross
Abstract: A large language model (LLM) can rate depression severity directly from a social media post or mark which clinical criteria the post shows and let co...
By Xinkai Chen
A large language model (LLM) can rate depression severity directly from a social media post or mark which clinical criteria the post shows and let code turn the count into a label. The latter is easie...
The authors present the Cross-Platform Fairness Evaluation (CPFE) framework, a five‑axis audit protocol that assesses discriminative performance, calibration, statistical significance, prediction equity, and attribution stability of transformer models. Applying CPFE to four models trained on a Kaggle mental‑health corpus and tested on Reddit and Twitter, they find substantial cross‑platform degradation in AUC (30–40%) and severe calibration failures (ECE rising to 0.5 on Twitter). The study demonstrates that platform‑specific temperature scaling can largely fix calibration without harming discrimination, while prediction equity and attribution stability analyses reveal significant disparities and vocabulary divergence across platforms. The results argue that cross‑platform validation across all CPFE axes should become a standard requirement for mental‑health NLP systems deployed in heterogeneous environments.
By Rajveer Singh Pall, Sameer Yadav
arXiv:2606. 27247v1 Announce Type: new Abstract: In NLP, mental health conditions are often modeled as isolated phenomena, without interpersonal context.
By Parmitha Vangapandu, Sai Ganesh Mokkapati, Sathwik Narkedimilli, MSVPJ Sathvik, Timothy Liu, Simon See, Johannes C. Eichstaedt
The paper introduces TSS (Triple-Stream Stress probe), a diagnostic framework that splits text into lexical, morpho-syntactic, and psycholinguistic style channels to analyze mental health NLP classifiers. Across four English datasets, TSS uncovers a lexical interference effect where adding lexical features harms performance on human-labeled data but not on auto-labeled data, and proposes the Degree of Divergence (DoD) statistic to audit label-source bias. The study demonstrates that style features largely remain effective even after masking content words, emphasizing that shortcut learning is label-source specific rather than clinically relevant.
By Moustafa Yehia Hassan
arXiv:2609.22767v1 Announce Type: cross
Abstract: The IEEE BigData Cup benchmark combines three prediction problems with different output structures: ordinal suicide-risk classification, multi-label...
By Zirui Li, Yanling Li, Kaolanglang Gao
The paper reports that in English all‑words word sense disambiguation (WSD), the scarcity of high‑quality labels—not the models—has become the limiting factor. The authors introduce lexEN, a human‑adjudicated correction layer over the Maru2022 ALL_NEW benchmark, and SenseBench, a living leaderboard for LLM WSD evaluation. They show that frontier large language models reach about 95 % accuracy on lexEN‑v1, that relabeling corpora with these models improves downstream systems, and that fine‑grained WordNet senses are often ill‑posed, with coarsening improving both annotator agreement and model performance.
"whyItMatters":"The study highlights that improving label quality and managing annotation costs are now the critical challenges for advancing WSD performance, as model accuracy is already near its theoretical ceiling."
By Vassili Philippov, Amro Salman, Dmitrii Andreev, Penny Hands, Emil Kaiumov, Pavel Katunin, Anton Nikolaev
arXiv:2608. 07316v1 Announce Type: cross Abstract: Natural Language Processing (NLP) models predicting mental health outcomes rarely specify what they measure: contextual knowledge, emotional content, or syntactic structure.
By Edoardo Sebastiano De Duro, Emma Franchino, Massimo Stella
The study examines how emotions are represented across layers of large language models (LLMs) by probing eight 1B–9B open‑weight models on three datasets (Twitter, Reddit, autobiographical narratives). It finds that the optimal probing layer varies systematically with the dataset, moving from near‑input layers to deeper layers, and that targeted forward‑pass interventions on these layers degrade performance more than random interventions. Additionally, the selected layers transfer across datasets and emotion categories, and early‑exit representations from these layers outperform full‑depth exits by an average of 6.9 percentage points.
By Tian Fang, Ga\"el Guibon, Davide Buscaldi
arXiv:2606. 00467v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used for zero-shot annotation and LLM-as-a-judge tasks, yet their reliability hinges on how model-internalized priors interact with user-provided instructions.
By Etienne Casanova, Rafal Kocielnik, R. Michael Alvarez
arXiv:2608. 07208v1 Announce Type: cross Abstract: Existing measures of how much a text is about a concept read the surface of the text: dictionary word shares, topic proportions, embedding similarities.
By Luc Hazenoot, Zhaochun Ren, Amirhossein Zohrehvand
arXiv:2607. 24817v1 Announce Type: cross Abstract: Digital mental health interventions (DMHIs) offer scalable support, but ensuring they accurately detect users' intent during volatile situations can be challenging.
By Anand Gupta, Akshat Surolia, Shubham Mishra, Shakil Imtiaz, Chaitali Sinha