arXiv AI

Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions

arXiv:2608. 13786v1 Announce Type: cross Abstract: Large language model (LLM) chatbots are increasingly used to answer clinical questions with citations to relevant clinical studies.

arXiv AI
Jun 6

Evaluating the Utility of Personal Health Records in Personalized Health AI

arXiv:2605. 18937v2 Announce Type: replace Abstract: Patient-managed Personal Health Records (PHRs) promises to empower patients to better understand their health; but information in the record is complex, potentially hindering insights.

By Rory Sayres, Kejia Chen, Ayush Jain, Matthew Thompson, Jonathan Richina, Xiang Yin, Jimmy Hu, Fan Zhang, Bob Lou, Mike Sanchez, Ines Mezerreg, Meredith Schreier, Hamsa Subramaniam, I-Ching Lee, Yugang Jia, Daniel Mcduff, Yossi Matias, Avinatan Hassidim, Dale Webster, Yun Liu, Jackie Barr, Quang Duong
arXiv AI
Jul 28

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

arXiv:2509. 02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions requires going beyond conventional benchmarks to assess how these systems behave in complex, high-stakes clinical scenarios.

By Sandhanakrishnan Ravichandran, Shivesh Kumar, Rogerio Corga Da Silva, Miguel Romano, Reinhard Berkels, Michiel van der Heijden, Olivier Fail, Valentine Emmanuel Gnanapragasam
arXiv AI
Sep 15

ClinAgent: A ReAct-Based Agent for Conversational Access to Clinical Trial Information

ClinAgent is a conversational system that uses a ReAct-based LLM agent to retrieve and synthesize clinical trial information from multiple sources such as ClinicalTrials.gov, PubMed, and a local dataset. The agent iteratively reasons over user queries, selects appropriate tools, and refines its actions to provide grounded, up-to-date responses in natural language across multi-turn interactions. Evaluation across three phases shows that DeepSeek (thinking mode) excels in planning quality while Gemini 3.0 Flash delivers the highest overall performance and expert ratings, demonstrating the promise of agentic AI for improving clinical trial data access.

By Antonino Vaccarella, Riccardo Cantini, Domenico Talia, Paolo Trunfio, Marianna Talia, Rosamaria Lappano, Marcello Maggiolini
arXiv AI
Sep 25

CRISS: A Retrieval-Augmented AI Chatbot for Assisting Cancer Registrars

CRISS is a retrieval‑augmented generation chatbot designed to aid cancer registrars by providing rapid, citation‑supported answers to complex coding and staging questions. It uses a domain‑specific knowledge base of national registry standards, indexed as dense embeddings, to retrieve relevant passages that ground responses generated by a large language model. In evaluations, RAG configurations consistently outperformed non‑RAG baselines across easy, medium, and hard questions, achieving higher grounding and semantic similarity scores while maintaining human oversight for final decisions.

By Vani Seth, Mohammad Beheshti, Anirudh Kambhampati, Vishwa Bhayani, Lucinda Ham, Prasad Calyam, Iris Zachary
arXiv Computation and Language
Aug 24

An ambiguity taxonomy for evaluating large language model performance on clinical registry abstraction: a multi-site prospective study

The study evaluates large language models (LLMs) on unprocessed electronic medical record data for clinical registry abstraction, focusing on the American College of Cardiology National Cardiovascular Data Registry. In a pilot at one academic center, the LLM identified candidate data sources for each registry question, which abstractors used to define question‑specific document sets. In a subsequent validation at a second center, the LLM answered 157 registry questions with an overall mean accuracy of 91.5%, but accuracy dropped from 96% for simple medication or event flag questions to 62% for event timing questions, reflecting increasing ambiguity and required clinical reasoning.

By James Matheson, Betsy Castillo, Andrew Y. Shin, David Scheinker
arXiv AI
Aug 13

A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench

arXiv:2608. 12138v1 Announce Type: cross Abstract: General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed largely in high-income settings.

By Praveen Reddy, Charuta Mandke, Suvrankar Datta, Sarah Khan, Siddharth Reddy Anthireddy, Shitij Arora, Vishal Singh
arXiv Computation and Language
Sep 10

A Patient Simulation Framework for Risk Assessment of Conversational Healthcare AI: Evaluation of an Antidepressant Decision Aid

arXiv:2602.11391v5 Announce Type: replace Abstract: Objective: This study develops and validates a patient simulation framework that aligns with the National Institute of Standards and Technology AI...

By Md Tanvir Rouf Shawon, Mohammad Sabik Irbaz, Hadeel R. A. Elyazori, Keerti Reddy Resapu, Yili Lin, Vladimir Franzuela Cardenas, K. Pierre Eklou, Farrokh Alemi, Kevin Lybarger
arXiv Computation and Language
Sep 4

Beyond Majority Vote: Multi-Perspective Adjudication for Medical Hallucination Detection

The paper presents a multi‑perspective annotation framework for detecting medical hallucinations in chatbot responses. It combines first‑pass annotators, a large language model acting as a judge (LaJ) for candidate discovery, and two adjudication stages—medical‑expert review and evidence‑based fact‑checking. The study finds that single‑pass labeling undercounts errors, while multi‑pass adjudication improves coverage but still depends on expert judgment and evidence.

By Joe Cecil, Marjorie Freedman
arXiv AI
2d ago

Scientific Agents: Evaluating Profession-Specific System Prompts on Scientific Tasks

The study evaluates whether detailed, profession‑specific system prompts improve performance on scientific tasks. Using an open‑source corpus of 503 agent profiles and Gemini 3.8 Flash, the authors compared matched profiles to four control prompts across nine text‑based science benchmarks and a tool‑using bioinformatics benchmark. Results show no consistent accuracy gains; matched profiles actually increased token usage and cost, and in some cases reduced success rates, with only a minor advantage in one benchmark likely due to prompt length rather than domain expertise.

By Timothy Kassis