arXiv Computation and Language
Aug 27

Retrieval-Augmented Agentic Rubric Generation for Reliable Medical Response Evaluation

The paper introduces a retrieval‑augmented multi‑agent framework that automatically generates instance‑specific evaluation rubrics for medical language models. By retrieving authoritative medical evidence, decomposing it into atomic facts, and combining these with user interaction constraints, the system produces fine‑grained criteria that outperform GPT‑4o on HealthBench and LLMEval‑Med. The generated rubrics also guide response refinement, improving medical LLM output quality by 9.2%.

By Yinzhu Chen, Abdine Maiga, Hossein A. Rahmani, Emine Yilmaz
arXiv AI
Aug 13

A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench

arXiv:2608. 12138v1 Announce Type: cross Abstract: General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed largely in high-income settings.

By Praveen Reddy, Charuta Mandke, Suvrankar Datta, Sarah Khan, Siddharth Reddy Anthireddy, Shitij Arora, Vishal Singh
arXiv AI
Sep 18

Guideline-grounded retrieval-augmented generation for ophthalmic clinical decision support

The paper introduces Oph‑Guid‑RAG, a multimodal retrieval‑augmented generation system tailored for ophthalmology clinical question answering. It treats each guideline page as an independent evidence unit, retrieving page images to preserve tables, flowcharts, and layout, and employs a controllable retrieval framework with routing and filtering to reduce noise. Evaluated on HealthBench, the system outperforms GPT‑5.2 and GPT‑5.4 on hard cases, achieving significant gains in overall score and accuracy, and ablation studies confirm the importance of reranking, routing, and retrieval design.

By Shuying Chen, Sen Cui, Zhong Cao
arXiv AI
Sep 25

A Living Benchmark for Information Retrieval from Electronic Health Records

The paper introduces BRIE, a continuously maintainable benchmark for evaluating large language models (LLMs) in electronic health record (EHR) information retrieval. It presents a scalable framework that automatically generates question–answer pairs from longitudinal EHR notes, validated by nineteen clinicians. The benchmark allows assessment of multiple inference strategies and highlights that state‑of‑the‑art LLMs often miss clinically important information, especially when synthesis across documents is required.

By Jordan L. Cahoon, Chloe O. Stanwyck, Sulaiman Somani, Philip Chung, Kevin R Keet, Kameron C. Black, Andrea T. Fisher, Sarita Khemani, Jerry Liu, Stephen Ma, Saloni K. Maharaj, Rita M. Pandya, Eduardo Perez-Guerrero, Priyanka Pillai, Lisa Shieh, David J. H. Wu, James Xie, James C. McAvoy, Teresa Nguyen, Jessica Tran, Lucy Yin, Bridget Lin, Alison Callahan, Jason A. Fries, Nigam H. Shah, Emily Alsentzer
arXiv AI
Sep 15

ClinAgent: A ReAct-Based Agent for Conversational Access to Clinical Trial Information

ClinAgent is a conversational system that uses a ReAct-based LLM agent to retrieve and synthesize clinical trial information from multiple sources such as ClinicalTrials.gov, PubMed, and a local dataset. The agent iteratively reasons over user queries, selects appropriate tools, and refines its actions to provide grounded, up-to-date responses in natural language across multi-turn interactions. Evaluation across three phases shows that DeepSeek (thinking mode) excels in planning quality while Gemini 3.0 Flash delivers the highest overall performance and expert ratings, demonstrating the promise of agentic AI for improving clinical trial data access.

By Antonino Vaccarella, Riccardo Cantini, Domenico Talia, Paolo Trunfio, Marianna Talia, Rosamaria Lappano, Marcello Maggiolini