DiaWhisper-DPO: Role-Attributed Transcription of Clinical Interviews via Failure-Mined Preference Optimization
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The paper introduces DiaWhisper-DPO, an end‑to‑end model that fine‑tunes Whisper-large-v3 with LoRA and a frame‑level role head to transcribe and attribute utterances in clinical interviews. It further refines the system using failure‑mined preference optimization (DPO) that leverages genuine decoding failures as rejected completions, eliminating the need for human preference data. On the DAIC‑WOZ dataset, DiaWhisper‑DPO attains 0.973 role accuracy and 0.119 DER, outperforming cascaded baselines by 72% and dramatically reducing seed variation, while also improving performance on the cross‑lingual PDCH‑HAMD dataset.
The paper examines automatic depression detection from doctor‑patient conversations and finds that models trained on semi‑structured interview data can achieve high accuracy by exploiting fixed interviewer prompts rather than the participants’ language. Across three datasets (ANDROIDS, DAIC‑WOZ, E‑DAIC), the authors show that restricting models to participant utterances distributes decision evidence more broadly and reflects genuine linguistic cues. The study highlights a cross‑dataset, architecture‑agnostic bias introduced by interviewer prompts and calls for analyses that localize decision evidence by time and speaker to ensure models learn from participants’ language.
arXiv:2609.38491v1 Announce Type: new Abstract: Clinical research in psychiatry increasingly relies on large scale collection of spoken language data to identify acoustic and linguistic biomarkers. Y...
arXiv:2609.14231v1 Announce Type: cross Abstract: Controllable synthesis of nonverbal vocalizations (NVVs) is es- sential for natural and expressive speech, but remains challeng- ing due to their aco...
arXiv:2609.28430v1 Announce Type: new Abstract: This work addresses continuous depression-severity score prediction from clinical interview transcripts under data scarcity. We propose a sequential lo...
Audio large language models (Audio LLMs) often fail to transcribe English‑Mandarin code‑switching speech, exhibiting language omission, translation‑instead‑of‑transcription, and hallucination. By applying Direct Preference Optimization (DPO) with 100 K preference pairs, the models learn to preserve mixed‑language content rather than translate, leading to significant reductions in mixed‑error rates (up to 89.6% in‑distribution). The study demonstrates that DPO can effectively align multilingual Audio LLMs for accurate code‑switching transcription.