Structure vs. Chain-of-Thought: Evaluating LLM Criteria Extraction for Depression Severity
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
A large language model (LLM) can rate depression severity directly from a social media post or mark which clinical criteria the post shows and let code turn the count into a label. The latter is easie...
The study compares six approaches—frontier commercial models, fine‑tuned smaller models, and conventional classifiers—for detecting anxiety in Reddit posts. It reveals a significant lexical bias: 69.3% of anxiety‑labelled posts contain the word "anxiety" or a variant, allowing models to perform well via keyword matching rather than true language understanding. After removing these terms, the frontier model still leads (F1 = 0.846), but a 110 M‑parameter domain‑adapted encoder achieves a close score (F1 = 0.831) without external API calls, and lexical dependence varies widely across models.
arXiv:2606. 05970v1 Announce Type: cross Abstract: Large language models are increasingly used for structured extraction from clinical free-text notes, but the sensitivity of their output to upstream configuration choices is less understood than their accuracy on fixed benchmarks.
The paper introduces TSS (Triple-Stream Stress probe), a diagnostic framework that splits text into lexical, morpho-syntactic, and psycholinguistic style channels to analyze mental health NLP classifiers. Across four English datasets, TSS uncovers a lexical interference effect where adding lexical features harms performance on human-labeled data but not on auto-labeled data, and proposes the Degree of Divergence (DoD) statistic to audit label-source bias. The study demonstrates that style features largely remain effective even after masking content words, emphasizing that shortcut learning is label-source specific rather than clinically relevant.
arXiv:2609.22767v1 Announce Type: cross Abstract: The IEEE BigData Cup benchmark combines three prediction problems with different output structures: ordinal suicide-risk classification, multi-label...
arXiv:2608. 08746v1 Announce Type: new Abstract: Prospective daily symptom tracking is central to premenstrual health assessment, but repeated ordinal forms impose substantial response burden.