Structure vs. Chain-of-Thought: Evaluating LLM Criteria Extraction for Depression Severity
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
arXiv:2609.39049v1 Announce Type: cross Abstract: A large language model (LLM) can rate depression severity directly from a social media post or mark which clinical criteria the post shows and let co...
arXiv:2606. 05970v1 Announce Type: cross Abstract: Large language models are increasingly used for structured extraction from clinical free-text notes, but the sensitivity of their output to upstream configuration choices is less understood than their accuracy on fixed benchmarks.
The study compares six approaches—frontier commercial models, fine‑tuned smaller models, and conventional classifiers—for detecting anxiety in Reddit posts. It reveals a significant lexical bias: 69.3% of anxiety‑labelled posts contain the word "anxiety" or a variant, allowing models to perform well via keyword matching rather than true language understanding. After removing these terms, the frontier model still leads (F1 = 0.846), but a 110 M‑parameter domain‑adapted encoder achieves a close score (F1 = 0.831) without external API calls, and lexical dependence varies widely across models.
arXiv:2608. 08746v1 Announce Type: new Abstract: Prospective daily symptom tracking is central to premenstrual health assessment, but repeated ordinal forms impose substantial response burden.
The paper introduces TSS (Triple-Stream Stress probe), a diagnostic framework that splits text into lexical, morpho-syntactic, and psycholinguistic style channels to analyze mental health NLP classifiers. Across four English datasets, TSS uncovers a lexical interference effect where adding lexical features harms performance on human-labeled data but not on auto-labeled data, and proposes the Degree of Divergence (DoD) statistic to audit label-source bias. The study demonstrates that style features largely remain effective even after masking content words, emphasizing that shortcut learning is label-source specific rather than clinically relevant.
The study investigates how a large language model, Gemma-3-27B-PT, internally represents depressive symptoms. By applying mechanistic interpretability methods to the model’s residual stream, researchers found that symptom groups are geometrically distinct at layer 21, and that projected symptom vectors align with clinician-annotated rankings across mood, somatic, and suicidality dimensions. Additionally, a single depression vector at this layer can differentiate depressive from non-depressive text with an AUC of 0.789, suggesting a potential emotional valence gate for symptom projection.