arXiv:2607. 06452v1 Announce Type: cross Abstract: Biomedical question answering requires not only accurate extraction of information from scientific literature but also reliable integration of evidence across multiple documents.
By Taeyun Roh, Eunha Lee, Wonjune Jang, Sohyun Chung, Junha Jung, Jaewoo Kang
MedRAGChecker is a claim-level verification framework designed for biomedical retrieval‑augmented generation (RAG). It decomposes generated answers into atomic claims and assesses each claim’s support by combining evidence‑grounded natural language inference with biomedical knowledge‑graph consistency signals. The aggregated claim decisions provide diagnostics that distinguish retrieval and generation failures, such as faithfulness, under‑evidence, contradiction, and safety‑critical errors, and the system is distilled into compact models for scalable evaluation.
By Yuelyu Ji, Min Gu Kwak, Hang Zhang, Xizhi Wu, Chenyu Li, Yanshan Wang
arXiv:2601.03471v4 Announce Type: replace-cross
Abstract: Reliable epidemiological reasoning requires synthesizing study evidence to infer disease burden, transmission dynamics, and intervention effe...
By Mingyang Wei, Dehai Min, Zewen Liu, Yuzhang Xie, Guanchen Wu, Ziyang Zhang, Carl Yang, Max S. Y. Lau, Qi He, Lu Cheng, Wei Jin
arXiv:2607. 19678v1 Announce Type: cross Abstract: AI-generated answers in high-stakes domains are often fluent but difficult to verify, especially when they contain multi-step reasoning rather than a single final answer.
By Guneet Singh Kohli, Yuxiang Zhou, Michael Sejr Schlichtkrull, Gregory E Dean, Maria Liakata
arXiv:2606. 04262v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for everyday health questions, including whether a user can safely take another dose of an over-the-counter (OTC) medication.
By Maroof Kousar, Yibo Hu
arXiv:2607. 12310v1 Announce Type: cross Abstract: While modern question answering (QA) systems excel on clean, schema-aligned corpora, real-world knowledge is rarely so neatly packaged.
By Michael Solodko, Steven Gong, Guangwei Yu, Satya Krishna Gorti, Jesse C. Cresswell, Victor Zhong
arXiv:2606. 15735v1 Announce Type: cross Abstract: Discharge summaries are crucial clinical documents containing the context of a patient's overall hospital stay, and are routinely reviewed by medical experts for patient readmission, ongoing care, and diagnostic decision-making.
By Jiyoun Kim, Muhan Yeo, Eunhye Jang, Jeewon Yang, Hangyul Yoon, Su Ji Lee, Hee Jo Han, Hee-Jae Jung, Doyun Kwon, Jun young Lee, Jaehun Lee, Jung-Oh Lee, Sunjun Kweon, Jong Hak Moon, Daseul Kim, Minjae Cho, Edward Choi
PetQA is a Korean long‑form question‑answering benchmark designed to assess veterinary knowledge and clinical reasoning in large language and vision‑language models. It comprises 10,076 text‑only and 8,751 multimodal QA pairs about dogs and cats, with expert veterinarian answers, and a test split called PetQA‑Bench that includes question type and clinical condition annotations. The study evaluates 18 models across zero‑shot, retrieval‑augmented generation, and supervised fine‑tuning settings using ROUGE, BERTScore, and LLM‑as‑a‑judge metrics, revealing current models’ strengths and limitations and underscoring the need for better adaptation methods for clinically reliable veterinary AI; translated versions in five languages are also provided.
By Taegyun Kim, Youngwook Ham, Jungwook Rhim, Ju-Hyun An, Sungkyu Park, Kunwoo Park
arXiv:2607. 11891v1 Announce Type: cross Abstract: The deployment of large language models (LLMs) in specialized domains like medical diagnostics and financial advisory necessitates evaluating capabilities beyond general knowledge.
By Megha Chakraborty, Darssan L. Eswaramoorthi, Het Riteshkumar Shah, Madhur Thareja, Michelle A Ihetu, Harshul Raj Surana, Kaushik Roy, Amit Sheth
The UIC-AIHealth4All system was presented for the ArchEHR-QA 2026 shared task on grounded question answering from electronic health records. It participated in evidence identification, answer generation, and answer‑evidence alignment, using an answer‑first pipeline that generates candidate answers with cited note sentences before classifying the full evidence set. The system ranked third in evidence identification, ninth in answer generation, and fifth in answer‑evidence alignment, and a linguistic analysis showed its outputs were harder to read than clinician‑authored references, highlighting the need for readability optimization in clinical NLP.
By Mohammad Arvan, Hossein Haeri, Natalie Parde, Rebecca T. Feinstein
arXiv:2510. 21084v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown strong potential for clinical decision support through their advanced language understanding and reasoning capabilities.
By Juntao Li, Haobin Yuan, Ling Luo, Yuanyuan Sun, Jian Wang, Hongfei Lin
arXiv:2604. 13201v2 Announce Type: replace-cross Abstract: Large language models are emerging as scientific assistants, but evaluating their ability to reason from empirical data remains challenging.
By Oliver Bentham, Vivek Srikumar