arXiv Computation and Language

Understanding Clinical Cognitive Dialogues Using Large Language Models

The paper introduces a de‑identified corpus of 33 in‑person cognitive assessment conversations, comprising 8,250 utterances annotated for three speaker roles and 56 dialogue acts. The authors benchmark large language models on fine‑grained dialogue‑act classification and next‑patient‑utterance generation, finding that instruction tuning and reasoning‑aware fine‑tuning improve performance but that models still struggle with closely related dialogue acts. The corpus and benchmark are presented as tools to measure interaction structure in cognitive assessments and to support future research on conversational markers, clinician education, and validated simulated patients.

arXiv Computation and Language
Aug 28

Evaluating Language Models in Realistic Conversational Contexts

The paper introduces UPHELD, a large benchmark of human-to-human dialogues written by professional script writers, featuring realistic turn densities and over 36,000 per-turn human annotations. It evaluates existing automatic metrics and LLM-as-a-judge methods, finding them unreliable against expert human judgment. Using UPHELD, the authors develop a Mixture-of-Judges framework that improves correlation with human assessments by about 30%.

By Ilija Subasic, Andrew Rabinovich, Zhao Chen
arXiv Computation and Language
Aug 27

MTDiag: A Multi-Turn Diagnostic Dataset Towards Clinically Meaningful LLM Evaluation

MTDiag is a newly released multi-turn diagnostic dialogue dataset designed to evaluate large language models (LLMs) in clinically meaningful ways. It is built from DDXPlus, MIMIC-IV, and AJCR case reports, covering both common emergency department presentations and rare conditions, and normalizes cases into a canonical schema using UMLS concept identifiers and ICD-10 codes. The dataset includes a UserLM‑8B utterance‑generation pipeline and physician‑validated natural‑language utterances, and introduces clinical knowledge‑grounded metrics that go beyond simple diagnostic accuracy for multi‑turn differential diagnosis tasks.

By Pia Chouayfati, Alexander M. Fichtl, Miriam Ansch\"utz, George Doumat, Georg Groh
arXiv Computation and Language
Aug 27

EduDial: Constructing a Large-scale Multi-turn Teacher-Student Dialogue Corpus

EduDial is a large-scale multi-turn teacher‑student dialogue corpus covering 345 core knowledge points and 34,250 dialogue sessions, designed around Bloom’s taxonomy and ten questioning strategies such as situational, ZPD, and metacognitive questioning. The dataset includes differentiated teaching strategies for students at varying cognitive levels to provide targeted guidance. Using EduDial, the authors trained EduDial‑LLM 32B and introduced an 11‑dimensional evaluation framework that measures teaching quality and content quality, showing that most mainstream LLMs struggle with student‑centered teaching while EduDial‑LLM outperforms all baselines across all metrics.

By Shouang Wei, Min Zhang, Xin Lin, Bo Jiang, Zhongxiang Dai, Kun Kuang
arXiv AI
Jul 15

Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking

arXiv:2607. 12085v1 Announce Type: new Abstract: Evaluating retail conversational agents requires methods beyond lexical-overlap metrics to assess intent alignment, factuality, helpfulness, clarity, tone, and overall response quality.

By Niranjan Kumar M, Balaji Nagarajan, Karthik Nair, Faysal Satter, Nithin Surendran
arXiv AI
Aug 28

When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue

The paper introduces ContraTalk, a benchmark that tests whether dialogue models truly use acoustic cues or rely on transcript shortcuts. It formalizes cross‑modal disagreement, creates conflict and consistent QA examples, and proposes an Audio Twin representation to expose acoustic evidence to models. Experiments show that while text‑only LLMs perform well on consistent cases, they falter on conflict cases, and AudioLLMs only partially mitigate this issue.

By Yen-Ju Lu, Yuzhe Wang, Yaohan Guan, Xiluo He, Jiarui Hai, Mingrui Liang, Kaavya Chaparala, Thomas Thebaud, Laureano Moro-Velazquez, Najim Dehak, Jesus Villalba