The paper presents a system for the IEEE BigData 2026 Cup on Explainable Suicide Risk Assessment on Social Media. It tackles three tasks—risk-level classification, evidence phrase extraction, and multi-label factor identification—using Qwen2.5-Instruct models adapted with quantized low-rank adaptation (QLoRA) and an answer-masked causal language-model objective. The final system achieved a composite score of 0.7738, with 0.8089 on Task 1 and 0.6919 on Task 2, demonstrating that task‑specific training and tailored aggregation improve performance across the three tasks.
By Xuan Zhong Feng, Geoffrey Martin, Hexin Dong, Yifan Peng
arXiv:2609.07766v1 Announce Type: cross
Abstract: Assessing suicide risk from social media text is a small-data, high-stakes setting requiring not only severity prediction but also supporting evidenc...
By Shlok Shelat, Shrey Salvi, Souvik Roy, Manas Gaur, Amit Sheth
The paper introduces Clinical Intent Extraction (CIE), a task that transforms fragmented clinical action annotations into complete structured records called Clinical Intent Representation (CIR). CIR decomposes each action into verb, type, coded target, timing, condition, request‑intent (aligned to HL7 FHIR) and modality, adding dimensions absent in prior datasets. By re‑expressing five heterogeneous corpora into CIR, the authors create CIRCA, a benchmark of 10,011 harmonized intents with human‑validated subsets, crosswalks, and a deterministic FHIR R4 mapper, and demonstrate that existing models perform poorly on the full task, highlighting the need for targeted development.
By Alexander Apartsin, Yehudit Aperstein
The paper introduces TSS (Triple-Stream Stress probe), a diagnostic framework that splits text into lexical, morpho-syntactic, and psycholinguistic style channels to analyze mental health NLP classifiers. Across four English datasets, TSS uncovers a lexical interference effect where adding lexical features harms performance on human-labeled data but not on auto-labeled data, and proposes the Degree of Divergence (DoD) statistic to audit label-source bias. The study demonstrates that style features largely remain effective even after masking content words, emphasizing that shortcut learning is label-source specific rather than clinically relevant.
By Moustafa Yehia Hassan
arXiv:2608. 03854v1 Announce Type: new Abstract: When decoder language models are used as classifiers, predicted class probabilities depend on implementation choices, including the prompt template, verbalizer (label-to-token mapping), and scoring rule, that are rarely treated as experimental variables.
By Anton Rasmussen, Hong Qin
The paper introduces a signed lexical gate that combines a sentence classifier’s logit margin with a sparse lexical model’s support for the predicted intent, assigning positive evidence to lexical agreement and negative evidence to a lexically favored competing intent. This gate retains more information than unsigned lexical confidence or a hard agreement rule and is calibrated via an independent binomial procedure to meet specified risk targets. Experiments on BANKING77, CLINC150, and HWU64 show that the proposed score reduces the area under the risk‑coverage curve by up to 15.8% and increases accepted coverage at low error rates, offering a compact, interpretable confidence enhancement for risk‑calibrated intent routing.
By Yezhou Cheng, Zehua Yang, Bojun Lin
The authors present the Cross-Platform Fairness Evaluation (CPFE) framework, a five‑axis audit protocol that assesses discriminative performance, calibration, statistical significance, prediction equity, and attribution stability of transformer models. Applying CPFE to four models trained on a Kaggle mental‑health corpus and tested on Reddit and Twitter, they find substantial cross‑platform degradation in AUC (30–40%) and severe calibration failures (ECE rising to 0.5 on Twitter). The study demonstrates that platform‑specific temperature scaling can largely fix calibration without harming discrimination, while prediction equity and attribution stability analyses reveal significant disparities and vocabulary divergence across platforms. The results argue that cross‑platform validation across all CPFE axes should become a standard requirement for mental‑health NLP systems deployed in heterogeneous environments.
By Rajveer Singh Pall, Sameer Yadav
arXiv:2607. 04223v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) reduces but does not eliminate hallucination, and existing detectors return a single answer-level score that does not indicate which sentence is unsupported, or why.
By Mohamed Aly Bouke
arXiv:2609.39049v1 Announce Type: cross
Abstract: A large language model (LLM) can rate depression severity directly from a social media post or mark which clinical criteria the post shows and let co...
By Xinkai Chen
arXiv:2608. 05162v1 Announce Type: cross Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden states into a passage-level vector, yet no shared protocol exists for comparing this choice across concepts, models, and tasks.
By Ayushi Agarwal
arXiv:2608. 14649v1 Announce Type: new Abstract: We present dLLM-SetScore, a training-free method that uses discrete masked-diffusion language models for multi-label text classification.
By Pawan Kumar
arXiv:2606. 27174v1 Announce Type: new Abstract: Medical device recalls are a critical regulatory mechanism for protecting patient safety.
By Ali Semih Atalay, Sevgi Yigit-Sert