arXiv AI

Condition-Gated Reasoning for Context-Dependent Biomedical Question Answering

arXiv:2602. 17911v3 Announce Type: replace-cross Abstract: Current biomedical question answering (QA) systems often assume that medical knowledge applies uniformly, yet real-world clinical reasoning is inherently conditional: nearly every decision depends on patient-specific factors such as comorbidities and contraindications.

arXiv AI
Sep 2

Multilingual Medical Reasoning for Question Answering with Large Language Models

The paper introduces a method for generating multilingual reasoning traces for medical question answering using large language models. It creates 500,000 reasoning traces in English, Italian, and Spanish by retrieving medical information from Wikipedia and applies them to MedQA and MedMCQA datasets extended into Italian and Spanish. The approach improves performance in both few‑shot in‑context learning and supervised fine‑tuning, achieving state‑of‑the‑art results for 8B‑parameter LLMs and releasing all resources for further research.

By Pietro Ferrazzi, Aitor Soroa, Rodrigo Agerri
Hugging Face Trending Papers
Sep 24

A Living Benchmark for Information Retrieval from Electronic Health Records

The paper introduces BRIE, a scalable framework that automatically creates question–answer pairs from longitudinal electronic health record notes, validated by nineteen clinicians. It offers a continuously maintainable benchmark for evaluating large language models in clinical settings, addressing limitations of manual, costly, and quickly outdated existing benchmarks. Experiments across nine LLMs and five inference strategies reveal that even state‑of‑the‑art systems often miss clinically important information, especially for synthesis‑heavy queries.

arXiv AI
Sep 25

A Living Benchmark for Information Retrieval from Electronic Health Records

The paper introduces BRIE, a continuously maintainable benchmark for evaluating large language models (LLMs) in electronic health record (EHR) information retrieval. It presents a scalable framework that automatically generates question–answer pairs from longitudinal EHR notes, validated by nineteen clinicians. The benchmark allows assessment of multiple inference strategies and highlights that state‑of‑the‑art LLMs often miss clinically important information, especially when synthesis across documents is required.

By Jordan L. Cahoon, Chloe O. Stanwyck, Sulaiman Somani, Philip Chung, Kevin R Keet, Kameron C. Black, Andrea T. Fisher, Sarita Khemani, Jerry Liu, Stephen Ma, Saloni K. Maharaj, Rita M. Pandya, Eduardo Perez-Guerrero, Priyanka Pillai, Lisa Shieh, David J. H. Wu, James Xie, James C. McAvoy, Teresa Nguyen, Jessica Tran, Lucy Yin, Bridget Lin, Alison Callahan, Jason A. Fries, Nigam H. Shah, Emily Alsentzer
arXiv Computation and Language
2d ago

BioMol-MQA: A Multi-Modal Question Answering Dataset For LLM Reasoning Over Bio-Molecular Interactions

BioMol-MQA is a new question‑answering dataset focused on polypharmacy that combines a multimodal knowledge graph—containing both text and molecular structure—with challenging questions designed to test large language models’ ability to retrieve and reason over this diverse information. The dataset highlights the limitations of current retrieval‑augmented generation systems, which typically handle only single‑modality text, by demonstrating that existing LLMs perform poorly unless provided with the necessary multimodal background data. This underscores the need for more robust RAG frameworks capable of integrating multiple data types for accurate responses.

By Saptarshi Sengupta, Shuhua Yang, Paul Kwong Yu, Fali Wang, Suhang Wang
arXiv AI
Jul 15

CANDI: Contextual Alignment for Niche Domains Question Answering

arXiv:2607. 11891v1 Announce Type: cross Abstract: The deployment of large language models (LLMs) in specialized domains like medical diagnostics and financial advisory necessitates evaluating capabilities beyond general knowledge.

By Megha Chakraborty, Darssan L. Eswaramoorthi, Het Riteshkumar Shah, Madhur Thareja, Michelle A Ihetu, Harshul Raj Surana, Kaushik Roy, Amit Sheth
arXiv AI
Jul 21

Enhancing LLMs' Clinical Reasoning with Real-World Data from a Nationwide Sepsis Registry

arXiv:2505. 02722v2 Announce Type: replace Abstract: Although large language models (LLMs) have demonstrated impressive reasoning capabilities across general domains, their effectiveness in real-world clinical practice remains limited.

By Junu Kim, Chaeeun Shim, Sungjin Park, Su Yeon Lee, Gee Young Suh, Chae-Man Lim, Seong Jin Choi, Song Mi Moon, Kyoung-Ho Song, Eu Suk Kim, Hong Bin Kim, Sejoong Kim, Chami Im, Dong-Wan Kang, Yong Soo Kim, Hee-Joon Bae, Sung Yoon Lim, Han-Gil Jeong, Edward Choi
arXiv Computation and Language
Aug 24

MedRAGChecker: Claim-Level Verification for Biomedical Retrieval-Augmented Generation

MedRAGChecker is a claim-level verification framework designed for biomedical retrieval‑augmented generation (RAG). It decomposes generated answers into atomic claims and assesses each claim’s support by combining evidence‑grounded natural language inference with biomedical knowledge‑graph consistency signals. The aggregated claim decisions provide diagnostics that distinguish retrieval and generation failures, such as faithfulness, under‑evidence, contradiction, and safety‑critical errors, and the system is distilled into compact models for scalable evaluation.

By Yuelyu Ji, Min Gu Kwak, Hang Zhang, Xizhi Wu, Chenyu Li, Yanshan Wang
arXiv AI
Jun 9

From Conflict to Consensus: Boosting Medical Reasoning via Multi-Round Agentic RAG

arXiv:2603. 03292v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) exhibit high reasoning capacity in medical question-answering, but their tendency to produce hallucinations and outdated knowledge poses critical risks in healthcare fields.

By Wenhao Wu, Zhentao Tang, Yafu Li, Shixiong Kai, Mingxuan Yuan, Zhenhong Sun, Chunlin Chen, Zhi Wang