arXiv AI

The complexities of patient-centred conversational artificial intelligence

arXiv:2607. 08625v1 Announce Type: new Abstract: Consumer-facing health chatbots powered by large language models (LLMs) are increasingly used for symptom assessment.

arXiv Computation and Language
Sep 1

SIC-Agents: Benchmarking and Building an Adaptive Simulator for Pediatric Serious Illness Communication Training

The paper introduces SIC-Agents, a self‑improving framework designed to enhance simulation for pediatric serious illness communication (SIC) training. It presents two new benchmark suites—PitfallBench and DialogueBench—that assess simulators at both turn‑level and full‑dialogue levels, specifically addressing the unique challenges of multi‑party interactions and parental distress. Experiments demonstrate that SIC‑Agents surpasses static expert prompting, and the authors release the benchmarks for broader research use.

By Zihan Wang, Anita Marie Slominska, Rennie Bimman, Elizabeth Di Flumeri, Amanda Mayappo-Neeposh, Conall Francoeur, Tamara Ellen Carver, Xiao-Wen Chang, Doina Precup, Esin Darici Haritaoglu, Ismail Haritaoglu, Akshatha Arodi, Naomi Goloff
arXiv AI
Aug 24

When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha

The paper evaluates the safety of conversational AI therapy bots for Generation Alpha, revealing that while these models understand 76‑82% of youth‑specific vocabulary, they correctly assess clinical risk only 64‑72% of the time, creating a significant vocabulary‑comprehension gap. Six failure patterns—such as sarcasm masking, minimization acceptance, and semantic drift—were identified, with compounded errors leading to a 94% miss rate when three or more patterns co‑occur. The authors estimate 146,880 missed crises annually and recommend mandatory human‑in‑the‑loop systems, quarterly youth‑specific validation, transparent performance disclosure, and regulatory oversight for youth‑facing mental health AI.

By Manisha Mehta, Virendra Mehta
arXiv Computation and Language
Sep 21

Generative Artificial Intelligence Chatbots for Motivational Interviewing: A Scoping Review From System Design to Intervention Outcomes

This scoping review examined 48 studies on generative AI chatbots designed to deliver motivational interviewing (MI). It found that most systems were text‑based and disembodied, with about half incorporating dynamic adaptation, and that safety reporting was inconsistent. While user perceptions were generally positive and many studies reported MI‑consistent interactions, evidence for sustained behavioral or functional change remains limited.

By Runze Hu, Jingqi Kong, Yang Yang, Yihang Yang, Jingyao Liu, Haizhou Tang, Shanghang Zhang, Zheng Liu
arXiv Computation and Language
Sep 10

A Patient Simulation Framework for Risk Assessment of Conversational Healthcare AI: Evaluation of an Antidepressant Decision Aid

arXiv:2602.11391v5 Announce Type: replace Abstract: Objective: This study develops and validates a patient simulation framework that aligns with the National Institute of Standards and Technology AI...

By Md Tanvir Rouf Shawon, Mohammad Sabik Irbaz, Hadeel R. A. Elyazori, Keerti Reddy Resapu, Yili Lin, Vladimir Franzuela Cardenas, K. Pierre Eklou, Farrokh Alemi, Kevin Lybarger
arXiv AI
Sep 15

Personalizing Personal Health Interfaces: Co-Design with Generative AI

The paper explores how generative AI can lower the barrier to personalizing health dashboards by enabling users to co-design interfaces in Figma Make. In a study with 14 participants, redesigns of Google and Apple Health focused on personal context, future planning, and interactive experiences, though conversational AI designs tended toward chat-window conventions. AI facilitated the materialization of loosely articulated ideas, yet model defaults and generation latency influenced iteration, and the process highlighted interpretability and accountability over privacy, trust, and emotional safety.

By Karthik S. Bhat, Vidhi Shah, Vedika Agnihotri, Dong Whi Yoo, Koustuv Saha
arXiv AI
Sep 15

K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations

arXiv:2609.15855v1 Announce Type: cross Abstract: % !TEX root = ../main.tex People increasingly use large language models (LLMs) for mental health support, yet their safety in evolving, high-risk con...

By Laura M. Vowels, Matthew J. Vowels, Shivali Sharma, Apoorv Jha, Rehnuma Choudhury, Wasseem El Sarraj, Rachel Francois-Walcott, Aruba Hussain, Sarah Ingram, Angela Loulopoulou, Adva Segal, Elena Volkova
arXiv AI
Jul 15

Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking

arXiv:2607. 12085v1 Announce Type: new Abstract: Evaluating retail conversational agents requires methods beyond lexical-overlap metrics to assess intent alignment, factuality, helpfulness, clarity, tone, and overall response quality.

By Niranjan Kumar M, Balaji Nagarajan, Karthik Nair, Faysal Satter, Nithin Surendran