arXiv:2608.22132v1 Announce Type: cross
Abstract: Biomedical multi-hop question answering (QA) requires models to connect evidence across intermediate entities such as diseases, drugs, proteins, and...
By Zhaohan Meng, Zaiqiao Meng, Siwei Liu, Hao Xu, Ke Yuan, Iadh Ounis
The paper introduces a method for generating multilingual reasoning traces for medical question answering using large language models. It creates 500,000 reasoning traces in English, Italian, and Spanish by retrieving medical information from Wikipedia and applies them to MedQA and MedMCQA datasets extended into Italian and Spanish. The approach improves performance in both few‑shot in‑context learning and supervised fine‑tuning, achieving state‑of‑the‑art results for 8B‑parameter LLMs and releasing all resources for further research.
By Pietro Ferrazzi, Aitor Soroa, Rodrigo Agerri
arXiv:2608.30556v1 Announce Type: new
Abstract: Path-finding over knowledge graphs has become an effective way to ground LLM reasoning on multi-hop questions. However, biomedical QA introduces two di...
By Jun Hyeong Kim, Dongki Kim, Yinhua Piao, Sung Ju Hwang
The paper introduces BRIE, a scalable framework that automatically creates question–answer pairs from longitudinal electronic health record notes, validated by nineteen clinicians. It offers a continuously maintainable benchmark for evaluating large language models in clinical settings, addressing limitations of manual, costly, and quickly outdated existing benchmarks. Experiments across nine LLMs and five inference strategies reveal that even state‑of‑the‑art systems often miss clinically important information, especially for synthesis‑heavy queries.
The paper introduces BRIE, a continuously maintainable benchmark for evaluating large language models (LLMs) in electronic health record (EHR) information retrieval. It presents a scalable framework that automatically generates question–answer pairs from longitudinal EHR notes, validated by nineteen clinicians. The benchmark allows assessment of multiple inference strategies and highlights that state‑of‑the‑art LLMs often miss clinically important information, especially when synthesis across documents is required.
By Jordan L. Cahoon, Chloe O. Stanwyck, Sulaiman Somani, Philip Chung, Kevin R Keet, Kameron C. Black, Andrea T. Fisher, Sarita Khemani, Jerry Liu, Stephen Ma, Saloni K. Maharaj, Rita M. Pandya, Eduardo Perez-Guerrero, Priyanka Pillai, Lisa Shieh, David J. H. Wu, James Xie, James C. McAvoy, Teresa Nguyen, Jessica Tran, Lucy Yin, Bridget Lin, Alison Callahan, Jason A. Fries, Nigam H. Shah, Emily Alsentzer
BioMol-MQA is a new question‑answering dataset focused on polypharmacy that combines a multimodal knowledge graph—containing both text and molecular structure—with challenging questions designed to test large language models’ ability to retrieve and reason over this diverse information. The dataset highlights the limitations of current retrieval‑augmented generation systems, which typically handle only single‑modality text, by demonstrating that existing LLMs perform poorly unless provided with the necessary multimodal background data. This underscores the need for more robust RAG frameworks capable of integrating multiple data types for accurate responses.
By Saptarshi Sengupta, Shuhua Yang, Paul Kwong Yu, Fali Wang, Suhang Wang