arXiv AI

FDARxBench: Benchmarking Regulatory and Clinical Reasoning on FDA Generic Drug Assessment

FDARxBench is an expert‑curated benchmark designed to evaluate document‑grounded question answering on FDA drug label documents, focusing on generic drug assessment. It features a multi‑stage pipeline that generates high‑quality QA examples covering factual, multi‑hop, and refusal tasks, and includes protocols for both open‑book and closed‑book reasoning. Experiments with various language models show significant gaps in factual grounding, long‑context retrieval, and safe refusal behavior, highlighting the challenge of regulatory‑grade label comprehension.

arXiv Computation and Language
Aug 24

MedRAGChecker: Claim-Level Verification for Biomedical Retrieval-Augmented Generation

MedRAGChecker is a claim-level verification framework designed for biomedical retrieval‑augmented generation (RAG). It decomposes generated answers into atomic claims and assesses each claim’s support by combining evidence‑grounded natural language inference with biomedical knowledge‑graph consistency signals. The aggregated claim decisions provide diagnostics that distinguish retrieval and generation failures, such as faithfulness, under‑evidence, contradiction, and safety‑critical errors, and the system is distilled into compact models for scalable evaluation.

By Yuelyu Ji, Min Gu Kwak, Hang Zhang, Xizhi Wu, Chenyu Li, Yanshan Wang
arXiv AI
Jun 16

EHRNote-ChatQA: A Benchmark for Evidence-Grounded Multi-Turn Clinical Question Answering over Longitudinal Discharge Summaries

arXiv:2606. 15735v1 Announce Type: cross Abstract: Discharge summaries are crucial clinical documents containing the context of a patient's overall hospital stay, and are routinely reviewed by medical experts for patient readmission, ongoing care, and diagnostic decision-making.

By Jiyoun Kim, Muhan Yeo, Eunhye Jang, Jeewon Yang, Hangyul Yoon, Su Ji Lee, Hee Jo Han, Hee-Jae Jung, Doyun Kwon, Jun young Lee, Jaehun Lee, Jung-Oh Lee, Sunjun Kweon, Jong Hak Moon, Daseul Kim, Minjae Cho, Edward Choi
arXiv AI
6d ago

PetQA: Benchmarking Veterinary Knowledge and Clinical Reasoning

PetQA is a Korean long‑form question‑answering benchmark designed to assess veterinary knowledge and clinical reasoning in large language and vision‑language models. It comprises 10,076 text‑only and 8,751 multimodal QA pairs about dogs and cats, with expert veterinarian answers, and a test split called PetQA‑Bench that includes question type and clinical condition annotations. The study evaluates 18 models across zero‑shot, retrieval‑augmented generation, and supervised fine‑tuning settings using ROUGE, BERTScore, and LLM‑as‑a‑judge metrics, revealing current models’ strengths and limitations and underscoring the need for better adaptation methods for clinically reliable veterinary AI; translated versions in five languages are also provided.

By Taegyun Kim, Youngwook Ham, Jungwook Rhim, Ju-Hyun An, Sungkyu Park, Kunwoo Park
arXiv AI
Jul 15

CANDI: Contextual Alignment for Niche Domains Question Answering

arXiv:2607. 11891v1 Announce Type: cross Abstract: The deployment of large language models (LLMs) in specialized domains like medical diagnostics and financial advisory necessitates evaluating capabilities beyond general knowledge.

By Megha Chakraborty, Darssan L. Eswaramoorthi, Het Riteshkumar Shah, Madhur Thareja, Michelle A Ihetu, Harshul Raj Surana, Kaushik Roy, Amit Sheth
arXiv Computation and Language
Aug 31

UIC-AIHealth4All at ArchEHR-QA 2026: Answer-First Evidence Grounding for Clinical Question Answering

The UIC-AIHealth4All system was presented for the ArchEHR-QA 2026 shared task on grounded question answering from electronic health records. It participated in evidence identification, answer generation, and answer‑evidence alignment, using an answer‑first pipeline that generates candidate answers with cited note sentences before classifying the full evidence set. The system ranked third in evidence identification, ninth in answer generation, and fifth in answer‑evidence alignment, and a linguistic analysis showed its outputs were harder to read than clinician‑authored references, highlighting the need for readability optimization in clinical NLP.

By Mohammad Arvan, Hossein Haeri, Natalie Parde, Rebecca T. Feinstein