arXiv AI

Natural Language Processing Psychometrics

arXiv:2608. 07316v1 Announce Type: cross Abstract: Natural Language Processing (NLP) models predicting mental health outcomes rarely specify what they measure: contextual knowledge, emotional content, or syntactic structure.

arXiv Machine Learning
Aug 28

Cross-Platform Generalisation Failure in Mental Health Natural Language Processing: A Five-Axis Fairness Audit of Transformer Models on Social Media

The authors present the Cross-Platform Fairness Evaluation (CPFE) framework, a five‑axis audit protocol that assesses discriminative performance, calibration, statistical significance, prediction equity, and attribution stability of transformer models. Applying CPFE to four models trained on a Kaggle mental‑health corpus and tested on Reddit and Twitter, they find substantial cross‑platform degradation in AUC (30–40%) and severe calibration failures (ECE rising to 0.5 on Twitter). The study demonstrates that platform‑specific temperature scaling can largely fix calibration without harming discrimination, while prediction equity and attribution stability analyses reveal significant disparities and vocabulary divergence across platforms. The results argue that cross‑platform validation across all CPFE axes should become a standard requirement for mental‑health NLP systems deployed in heterogeneous environments.

By Rajveer Singh Pall, Sameer Yadav
arXiv AI
Sep 3

Interpretable Symptom Vectors for Depression in a Large Language Model

The study investigates how a large language model, Gemma-3-27B-PT, internally represents depressive symptoms. By applying mechanistic interpretability methods to the model’s residual stream, researchers found that symptom groups are geometrically distinct at layer 21, and that projected symptom vectors align with clinician-annotated rankings across mood, somatic, and suicidality dimensions. Additionally, a single depression vector at this layer can differentiate depressive from non-depressive text with an AUC of 0.789, suggesting a potential emotional valence gate for symptom projection.

By Fangyi Zhu, Ajay Subramanian, Allison Constant, Camille Wang, Ravish Gupta, Corey J. Keller
arXiv AI
Aug 21

Computational Phenomenology of Borderline Personality Disorder: A Comparative Evaluation of LLM-Simulated Expert Personas and Human Clinical Experts

arXiv:2508. 19008v3 Announce Type: replace Abstract: Building on a human-led thematic analysis of clinical life-story interviews (> 150,000 words) with inpatients with Borderline Personality Disorder, this study examines the capacity of large language models (OpenAI's GPT, Google's Gemini, and Anthropic's Claude) to support qualitative clinical analysis.

By Marcin Moskalewicz, Anna Sterna, Karolina Dro\.zd\.z, Kacper Dudzic, Marek Pokropski, Paula Flores
arXiv Machine Learning
Aug 31

Multilingual Lexical Feature Analysis of Spoken Language for Predicting Major Depression Symptom Severity

arXiv:2511.07011v2 Announce Type: replace-cross Abstract: Background: Remotely captured spoken language could provide objective, regular indicators of depression symptom severity. However, research t...

By Anastasiia Tokareva, Judith Dineley, Zoe Firth, Pauline Conde, Faith Matcham, Sara Siddi, Femke Lamers, Ewan Carr, Carolin Oetzmann, Daniel Leightley, Yuezhou Zhang, Amos A. Folarin, Josep Maria Haro, Brenda W. J. H. Penninx, Raquel Bailon, Srinivasan Vairavan, Til Wykes, Richard J. B. Dobson, Vaibhav A. Narayan, Matthew Hotopf, Nicholas Cummins, The RADAR-CNS Consortium
arXiv Computation and Language
Sep 21

Reading Anxiety or Reading the Label? Comparing Fine-Tuned and Frontier Models for Anxiety Detection on Social Media

The study compares six approaches—frontier commercial models, fine‑tuned smaller models, and conventional classifiers—for detecting anxiety in Reddit posts. It reveals a significant lexical bias: 69.3% of anxiety‑labelled posts contain the word "anxiety" or a variant, allowing models to perform well via keyword matching rather than true language understanding. After removing these terms, the frontier model still leads (F1 = 0.846), but a 110 M‑parameter domain‑adapted encoder achieves a close score (F1 = 0.831) without external API calls, and lexical dependence varies widely across models.

By Cris Huynh, Arlene Pham
arXiv AI
Aug 24

The Divergence Hypothesis: Unmasking Lexical Interference and Label Bias in Mental Health NLP

The paper introduces TSS (Triple-Stream Stress probe), a diagnostic framework that splits text into lexical, morpho-syntactic, and psycholinguistic style channels to analyze mental health NLP classifiers. Across four English datasets, TSS uncovers a lexical interference effect where adding lexical features harms performance on human-labeled data but not on auto-labeled data, and proposes the Degree of Divergence (DoD) statistic to audit label-source bias. The study demonstrates that style features largely remain effective even after masking content words, emphasizing that shortcut learning is label-source specific rather than clinically relevant.

By Moustafa Yehia Hassan
arXiv AI
Sep 10

A Three-Tier Persona Vector for Controllable User Simulation in Agentic Evaluation

The paper introduces a three-tier persona vector for user simulation in evaluating LLM agents, comprising 23 dimensions across demographics, behavioral traits, and emotional states, plus a query-complexity overlay. It demonstrates that these nuanced personas generate diverse, scenario-reactive conversations, leading to significant variations in agent goal achievement and compliance across different contexts. The model’s design allows for reproducible, auditable user behavior patterns without relying on learned covariance matrices.

By Rahul Khedar, Eshita, Sneha Teja Sree Reddy Thondapu, Mayank Malhotra, Arup Kumar Das, Jitesh Chandra Mishra, Arun Menon, Avinash Karn, Mouli V
arXiv Computation and Language
Sep 10

A Patient Simulation Framework for Risk Assessment of Conversational Healthcare AI: Evaluation of an Antidepressant Decision Aid

arXiv:2602.11391v5 Announce Type: replace Abstract: Objective: This study develops and validates a patient simulation framework that aligns with the National Institute of Standards and Technology AI...

By Md Tanvir Rouf Shawon, Mohammad Sabik Irbaz, Hadeel R. A. Elyazori, Keerti Reddy Resapu, Yili Lin, Vladimir Franzuela Cardenas, K. Pierre Eklou, Farrokh Alemi, Kevin Lybarger