MedRAGChecker is a claim-level verification framework designed for biomedical retrieval‑augmented generation (RAG). It decomposes generated answers into atomic claims and assesses each claim’s support by combining evidence‑grounded natural language inference with biomedical knowledge‑graph consistency signals. The aggregated claim decisions provide diagnostics that distinguish retrieval and generation failures, such as faithfulness, under‑evidence, contradiction, and safety‑critical errors, and the system is distilled into compact models for scalable evaluation.
By Yuelyu Ji, Min Gu Kwak, Hang Zhang, Xizhi Wu, Chenyu Li, Yanshan Wang
arXiv:2603. 05308v3 Announce Type: replace-cross Abstract: Assessing whether an article supports an assertion is essential for hallucination detection and claim verification.
By Qiao Jin, Yin Fang, Lauren He, Yifan Yang, Guangzhi Xiong, Zhizheng Wang, Nicholas Wan, Joey Chan, Donald C. Comeau, Robert Leaman, Charalampos S. Floudas, Aidong Zhang, Michael F. Chiang, Yifan Peng, Zhiyong Lu
arXiv:2608.29582v1 Announce Type: cross
Abstract: Current evaluations of large language models (LLMs) primarily focus on factual knowledge retrieval, overlooking the fundamental challenge of navigati...
By Yi Yu, Bo Wang, Chong Feng, Ge Shi, Xia Liu, Ziyi Yang, Xuewen Shi
The article presents a new semantic model for representing scientific evidence, specifically tailored to genetics, that extends existing standards by adding fine‑grained, domain‑specific structure. It aligns with FHIR Evidence and SEPIO, incorporates a compact vocabulary validated by SHACL, and was tested in a human‑AI annotation pilot on six genetics papers, producing 28 evidence items and 95 source‑anchored assertions. The authors argue that this model advances trustworthy, AI‑ready infrastructure for variant interpretation by providing a reference data model and validation schema for genetic evidence.
By Michael Bouzinier, Dmitry Etin
The article proposes a framework called quantitative evidence mining to transform biomedical findings into structured, context-rich evidence units. It outlines core elements such as claim, measured entity, value, comparator, population, conditions, temporal context, uncertainty, provenance, validation, and expert review. The authors present an eight-stage reference architecture and emphasize that plausibility should remain multidimensional rather than collapsed into a single truth label, linking extraction to evidence synthesis for applications like clinical trials, biomarker research, and knowledge-graph construction.
By Negin Sadat Babaiha, Stefan Geissler, Marie-Christine Simon, Martin Hofmann-Apitius, Marc Jacobs
arXiv:2608.28607v1 Announce Type: new
Abstract: Pharmaceutical sponsors developing a drug for both the United States and the European Union must reconcile guidance issued independently by the FDA and...
By Chuchu Wu, Zhiyin Zhou, Jingzhuo Hu, Liang You
The paper investigates why retrieval‑based open‑ended evaluation fails in medical fact verification. By creating two detailed taxonomies—one for retrieval‑stage errors across five quality dimensions and another for verifier‑reasoning errors across six steps—the authors automatically label evidence quality and reasoning errors using an LLM‑as‑Judge pipeline. Their large‑scale stress tests across multiple retrieval methods and verifier models show that increasing model size, reasoning effort, source breadth, or medical fine‑tuning does not eliminate these failure modes, indicating fundamental limits of the retrieve‑then‑verify paradigm in open‑ended medical contexts.
By Heyuan Huang, Jirui Dai, Alexandra DeLucia, Sonal Joshi, Mahsa Yarmohammadi, Jie Gao, Bernal Jim\'enez Guti\'errez, Mark Dredze
The paper introduces a comparative explainability framework for auditing DeBERTa‑v3 in zero‑shot medical abstract classification. It evaluates five explanation methods—SHAP, LIME, occlusion, Input × Gradient, and Attention × Gradient—using a natural language inference engine on a balanced corpus of 1,000 abstracts per diagnostic category. The study finds that explanatory stability aligns with predictive certainty, identifies three systemic failure mechanisms, and recommends combining multiple explanation methods and quantitative agreement metrics for transformer‑based medical text classifiers.
By Javier Diaz Esteban-Herreros, David Mu\~noz-Valero, Raquel Mart\'inez-Espa\~na, Jose M. Juarez, Juan Moreno-Garcia
arXiv:2606.22419v3 Announce Type: replace
Abstract: A recent Nature Medicine study reports that general-purpose frontier LLMs outperform specialized retrieval-augmented clinical tools on medical benc...
By Madhulatha Mandarapu, Sandeep Kunkunuru
The paper introduces Distilled Rapid Embedding Transfer (DRET), a parameter‑efficient method that injects biomedical domain knowledge from large specialized models into a smaller general‑purpose model without retraining on the original specialized corpora. DRET evolves through iterative strategies—tokenizer‑merge (DRET 1.x), hybrid embedding averaging (DRET 2.0), priority‑based embedding transfer (DRET 3.x), and further refinements (DRET 4.x)—and demonstrates that a 66‑million‑parameter DistilBERT can achieve competitive or superior performance on token‑level PICO classification compared to much larger models, while remaining lightweight. The authors validate the embedding‑level transfer with cosine similarity, semantic‑shift, and t‑SNE analyses, highlighting DRET’s potential for scalable, resource‑efficient biomedical text mining.
By Girish Sundaram, Daniel Berleant
arXiv:2601.03418v3 Announce Type: replace
Abstract: Trustworthy clinical summarization requires every claim to be traceable to its evidence, yet existing attribution often resolves only to the senten...
By Bohao Chu, Hendrik Damm, Tabea M. G. Pakull, Sameh Frihat, Georg Lodde, Elisabeth Livingstone, Christoph M. Friedrich, Norbert Fuhr
arXiv:2606. 07141v1 Announce Type: cross Abstract: Language models trained for clinical disease inference are trained on patient data, which may include sensitive and private information, and data owners may request the removal of their data from a trained model due to privacy or copyright concerns.
By Anurag Sharma, Sai Teja Chunchu, Prasenjit Mitra, Sandipan Sikdar, Koustav Rudra