arXiv:2603. 05308v3 Announce Type: replace-cross Abstract: Assessing whether an article supports an assertion is essential for hallucination detection and claim verification.
By Qiao Jin, Yin Fang, Lauren He, Yifan Yang, Guangzhi Xiong, Zhizheng Wang, Nicholas Wan, Joey Chan, Donald C. Comeau, Robert Leaman, Charalampos S. Floudas, Aidong Zhang, Michael F. Chiang, Yifan Peng, Zhiyong Lu
The article presents a new semantic model for representing scientific evidence, specifically tailored to genetics, that extends existing standards by adding fine‑grained, domain‑specific structure. It aligns with FHIR Evidence and SEPIO, incorporates a compact vocabulary validated by SHACL, and was tested in a human‑AI annotation pilot on six genetics papers, producing 28 evidence items and 95 source‑anchored assertions. The authors argue that this model advances trustworthy, AI‑ready infrastructure for variant interpretation by providing a reference data model and validation schema for genetic evidence.
By Michael Bouzinier, Dmitry Etin
OpenMTB‑Audit is an open‑source benchmark that tests large language models on 500 synthetic non‑small cell lung cancer cases, covering five adversarial error categories and four safety labels: Supported, Partially Supported, Unsupported, and Insufficient Information. The study found that all eight tested LLMs over‑refused Partially Supported recommendations, collapsing labels to achieve high safety scores. A deterministic seven‑module framework, MTB‑AuditAgent, was introduced to reduce over‑refusal to 6.7% and reach 91.2% accuracy, while an oncologist annotation study highlighted disagreement around the boundary between information sufficiency and treatment optimization.
By Negin Ashrafi, Jia Luo, Stacey M. Frumm, Roxana Daneshjou
arXiv:2608. 02615v1 Announce Type: cross Abstract: Cancer diagnosis and characterization require integrating complementary evidence from radiology, pathology, genomics, and clinical metadata.
By Ahnaf Munir, Dannong Wang, Michael W. McDonald, Mubarak Shah, Pegah Khosravi, Yu Tian
The article proposes a framework called quantitative evidence mining to transform biomedical findings into structured, context-rich evidence units. It outlines core elements such as claim, measured entity, value, comparator, population, conditions, temporal context, uncertainty, provenance, validation, and expert review. The authors present an eight-stage reference architecture and emphasize that plausibility should remain multidimensional rather than collapsed into a single truth label, linking extraction to evidence synthesis for applications like clinical trials, biomarker research, and knowledge-graph construction.
By Negin Sadat Babaiha, Stefan Geissler, Marie-Christine Simon, Martin Hofmann-Apitius, Marc Jacobs
arXiv:2608.30393v1 Announce Type: new
Abstract: Biomedical artificial intelligence (AI) systems increasingly extract, organize, and reuse scientific claims from literature, clinical trials, and regul...
By Negin Sadat Babaiha, Stefan Geissler, Marie-Christine Simon, Martin Hofmann-Apitius, Marc Jacobs
arXiv:2610.01616v1 Announce Type: cross
Abstract: The emergence of foundation models for molecular property prediction requires a high degree of AI data readiness, including reliable metadata annotat...
By Laura van Weesep, Riccardo Tedoldi, Jens Sj\"olund, Hossein Azizpour, Susanne Winiwarter, Ola Engkvist, Jon Paul Janet, Samuel Genheden, Juan Viguera Diez
arXiv:2606. 19345v1 Announce Type: cross Abstract: The rapid increase in scientific publications leads to the fact that manual study screening in systematic literature reviews (SLRs) is increasingly resource consuming, inefficient, and inconsistent.
By Zhyar Rzgar K. Rostam, M\'arta P\'entek, J\'anos Tibor Czere, Zsombor Zrubka, L\'aszl\'o Gul\'acsi, G\'abor Kert\'esz
arXiv:2606. 11830v1 Announce Type: new Abstract: Background.
By Qianyu Yao, Fei Sun, Bocheng Huang, Wei Chen, Jiarui Jiang, Shu Quan, Yifei Chen, Wenjie Xu, Bo li, Liping Su, Ruoqiong Wu, Huhai Hong, Huimei Wang
MedFabric is a new benchmark for detecting word‑level medical fabrications, comprising 646 fabricated statements each paired with a ground‑truth passage that shares the same LLM authorship and nearly identical wording. The study shows that current detectors perform poorly—expert clinicians achieve only 53.3% macro‑F1 and no detector family surpasses 60% without gold evidence—highlighting that detection hinges on evidence correctness rather than subtlety of fabrication. The authors demonstrate that a retrieval‑confidence gate can substantially improve performance, raising macro‑F1 from 61% to 74%.
By Tung Sum Thomas Kwok, Qian Qian, Xiaofeng Lin, Dongxu Zhang, Jun Han, Zhichao Yang, Davin Hill, Tamer Soliman, Sanjit Singh Batra, Robert Tillman, Guang Cheng
arXiv:2601. 12805v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown growing promise in biomedical research, particularly for knowledge-driven interpretation tasks.
By Xiaohan Huang, Meng Xiao, Chuan Qin, Qingqing Long, Jinmiao Chen, Yuanchun Zhou, Hengshu Zhu
arXiv:2608.28974v1 Announce Type: new
Abstract: Clinically relevant oncology information is distributed across heterogeneous, longitudinal documentation, creating substantial abstraction burden and r...
By Daniel Kang, Michelle Hu, Soorya Ram Shimgekar, Shayan Vassef, Yufan Wang, Anit Kumar Sahu, Munmun De Choudhury, Vedant Das Swain, Christian Poellabauer, Li Yan Khor, Koustuv Saha, Robert Wojciechowski, Elliot Kidd, Piyum Zonooz, Navin Kumar