arXiv Machine Learning
Sep 21

How Many Humans Is a Judge Panel Worth?

arXiv:2609.21277v1 Announce Type: cross Abstract: How many human judgments does a panel of language models represent? The answer depends on what is matched. We audit categorical judge panels against...

By Chao Li, Yingying Yu, Yunfeng Li
arXiv AI
Aug 24

The Divergence Hypothesis: Unmasking Lexical Interference and Label Bias in Mental Health NLP

The paper introduces TSS (Triple-Stream Stress probe), a diagnostic framework that splits text into lexical, morpho-syntactic, and psycholinguistic style channels to analyze mental health NLP classifiers. Across four English datasets, TSS uncovers a lexical interference effect where adding lexical features harms performance on human-labeled data but not on auto-labeled data, and proposes the Degree of Divergence (DoD) statistic to audit label-source bias. The study demonstrates that style features largely remain effective even after masking content words, emphasizing that shortcut learning is label-source specific rather than clinically relevant.

By Moustafa Yehia Hassan