arXiv Computation and Language

Cross-Scale Transfer Learning for Depression Severity Prediction: From PHQ-8 to HAMD-17 Across Languages and Clinical Paradigms

arXiv Computation and Language
Sep 16

DiaWhisper-DPO: Role-Attributed Transcription of Clinical Interviews via Failure-Mined Preference Optimization

The paper introduces DiaWhisper-DPO, an end‑to‑end model that fine‑tunes Whisper-large-v3 with LoRA and a frame‑level role head to transcribe and attribute utterances in clinical interviews. It further refines the system using failure‑mined preference optimization (DPO) that leverages genuine decoding failures as rejected completions, eliminating the need for human preference data. On the DAIC‑WOZ dataset, DiaWhisper‑DPO attains 0.973 role accuracy and 0.119 DER, outperforming cascaded baselines by 72% and dramatically reducing seed variation, while also improving performance on the cross‑lingual PDCH‑HAMD dataset.

By Weiming Li, Ana Catarina Fidalgo Barata, Miguel Constante, Jo\~ao Miguel Sanches
arXiv AI
Sep 4

LongCounsel-8: A Benchmark Suite for Longitudinal Depression Tracking from Multi-Session Counseling Dialogues

LongCounsel-8 is a new benchmark suite comprising three datasets with 7,749 five‑session counseling dialogues, each grounded in real client profiles, depression trajectories, symptom compositions, and counseling patterns. The benchmark addresses challenges of longitudinal consistency, empirical grounding of symptom progression, and natural expression of controlled depression states without exposing labels. Experiments show that lower single‑session error does not ensure accurate trend detection, methods perform worse on worsening trajectories, and adding more session history can reduce trend prediction accuracy.

By Jiayi Li, Zhaomin Wu, Bingsheng He
arXiv Computation and Language
4d ago

TriageRA-CCF: Source-Side Clinical Confidence and Coverage Signals for Adaptive Rank Budgeting in Medical LLMs

The paper introduces TriageRA-CCF, a method for adaptively allocating low‑rank LoRA channels in medical large language models based on source‑side signals: answer confidence, clinical coverage, and a counterfactual close‑miss proxy. By supervising a budget router that selects among 2, 4, or 8 active ranks, the approach improves average accuracy over existing LoRA variants on Qwen3‑8B and Llama3.1‑8B, though gains vary across benchmarks. Ablation studies confirm that each signal contributes to better budget decisions, though their combined effect is not uniformly superior across all backbones.

By Shucan Ji, Yining Huang, Hongliang Guo
arXiv AI
Aug 24

Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization

The paper critiques the common practice of evaluating specialist language models solely by endpoint gains, arguing that the underlying model updates are largely unexamined. It introduces a paired weight‑delta path audit and applies it to two public generalist‑to‑medical‑specialist checkpoint pairs, showing that the decoder‑side updates largely reconstruct the observed benchmark improvements. However, the audit finds that the improvements are not neatly localized to a single component family, indicating that component‑level explanations are more complex than previously assumed.

By Praphul Singh, Shanu Kumar, Akshat Agarwal
arXiv Machine Learning
Jul 31

Psych-ECA: A Reproducible Semi-Synthetic Benchmark for Synthetic Control Arms in Longitudinal Psychiatry

arXiv:2607. 27224v1 Announce Type: cross Abstract: External and synthetic control arms (ECAs) are entering psychiatric drug development, but the field lacks a benchmark that evaluates the properties regulators care about: not only how accurately a method reconstructs untreated trajectories, but whether its uncertainty is calibrated, whether it is robust to the informative observation times common in mental-health records (sicker patients are seen more often), and what false-positive rate it induces in go/no-go trial decisions.

By Aakash Bhagat, Shashank Choudhary
arXiv AI
Jun 30

Primary ICD Category Prediction using LLM-based Probing

arXiv:2606. 28798v1 Announce Type: new Abstract: Objective: ICD codes are central to reimbursement, research, and population health surveillance, yet automated coding systems often struggle to integrate diagnostic signals from both clinical narratives and structured electronic health record (EHR) variables.

By Chengyuan Liu, Xinyue Zhang, Yao Li, Guanting Chen