arXiv:2608.28592v1 Announce Type: new
Abstract: Large language models (LLMs) achieve high scores on medical knowledge examinations, yet real-world oncology is not a knowledge test--it is a sequence o...
By Zhang Sheng, Jinming Li, Wangyang Chen, Zhiwei Bao, Yu YoSean Wang
OpenMTB‑Audit is an open‑source benchmark that tests large language models on 500 synthetic non‑small cell lung cancer cases, covering five adversarial error categories and four safety labels: Supported, Partially Supported, Unsupported, and Insufficient Information. The study found that all eight tested LLMs over‑refused Partially Supported recommendations, collapsing labels to achieve high safety scores. A deterministic seven‑module framework, MTB‑AuditAgent, was introduced to reduce over‑refusal to 6.7% and reach 91.2% accuracy, while an oncologist annotation study highlighted disagreement around the boundary between information sufficiency and treatment optimization.
By Negin Ashrafi, Jia Luo, Stacey M. Frumm, Roxana Daneshjou
arXiv:2609.38181v1 Announce Type: new
Abstract: Survival analysis estimates time-to-event outcomes from patient covariates and is widely used for medical risk assessment. Patients seeking prognostic...
By Juan M Zambrano Chaves, Peniel Argaw, Risa Ueno, Carlo Bifulco, Kristina Young, Rom Leidner, Tristan Naumann, Hoifung Poon
arXiv:2510. 17532v2 Announce Type: replace-cross Abstract: Predicting cancer treatment outcomes requires models that are both accurate and interpretable, particularly in the presence of heterogeneous clinical data.
By Raghu Vamshi Hemadri, Geetha Krishna Guruju, Kristi Topollai, Anna Ewa Choromanska
arXiv:2608. 16594v1 Announce Type: new Abstract: Cancer survival prediction supports treatment planning, risk stratification, and follow-up management.
By Tianqi Xiang, Qixiang Zhang, Xinpeng Ding, Yi Li, Xiaomeng Li
arXiv:2606. 19183v1 Announce Type: cross Abstract: Large language models (LLMs) can make clinical decision support more accessible by interpreting free-text documentation, but their direct use as diagnostic engines is limited by sensitivity to prompts, information order, and plausible but incorrect outputs.
By Soheyl Bateni, Maryam Abdolali
arXiv:2606. 06224v1 Announce Type: cross Abstract: Explanations of multiple instance learning (MIL) models are widely used for validation and discovery in digital histopathology.
By Yanqing Luo (Berlin Institute for the Foundations of Learning and Data, Berlin, Germany, Machine Learning Group, Technische Universit\"at Berlin, Berlin, Germany), Julius Hense (Berlin Institute for the Foundations of Learning and Data, Berlin, Germany, Machine Learning Group, Technische Universit\"at Berlin, Berlin, Germany), Niklas Preni{\ss}l (Institute of Pathology, Charit\'e Universit\"atsmedizin, Berlin, Germany, Berlin Institute of Health at Charit\'e -- Universit\"atsmedizin Berlin, BIH Biomedical Innovation Academy, BIH Charit\'e Digital Clinician Scientist Program, Berlin, Germany), Andreas Mock (Institute of Pathology, Ludwig Maximilian University of Munich, Munich, Germany, Division of Translational Medical Oncology, DKFZ, Heidelberg, Germany, NCT Heidelberg, Heidelberg, Germany, German Cancer Consortium), Klaus-Robert M\"uller (Berlin Institute for the Foundations of Learning and Data, Berlin, Germany, Machine Learning Group, Technische Universit\"at Berlin, Berlin, Germany, Department of Artificial Intelligence, Korea University, Seoul, Korea, Max-Planck Institute for Informatics, Saarbr\"ucken, Germany), Thomas Schnake (Department of Chemistry, Chemical Physics Theory Group, University of Toronto, Canada, Vector Institute for Artificial Intelligence, Toronto, Canada, Acceleration Consortium, University of Toronto, Canada), Mina Jamshidi Idaji (Berlin Institute for the Foundations of Learning and Data, Berlin, Germany, Machine Learning Group, Technische Universit\"at Berlin, Berlin, Germany)
arXiv:2608. 02803v1 Announce Type: cross Abstract: Attention-based multiple instance learning (ABMIL) is the predominant approach for slide-level prediction in computational pathology, yet its attention maps provide only local explanations: they indicate where a model focuses but not which histological features drive its predictions or how the model behaves across a patient cohort.
By Abdallah Lamane, Abdul Rahman Diab, Ren-Chin Wu, William Lotter
arXiv:2609.01202v1 Announce Type: cross
Abstract: Clinical trials are essential for advancing cancer care and drug development, but many fail because of insufficient patient enrollment. While there i...
By Yin Fang, Qiao Jin, Shubo Tian, Lauren He, Maya Geer, Noor Naffakh, Ryan Huu-Tuan Nguyen, Zifeng Wang, Jimeng Sun, Charalampos S. Floudas, James L. Gulley, Kamilia Moalem, Catarina Martins Maia, Amanda Nottke, Juan W. Valle, Melinda Bachini, Lourdes Rocha-Nussbaum, Kari Ramage, Nikita Curry, Megan Barnes, Mandy Mansaray, Darlene Gabeau, Craig E. Grossman, Heath Skinner, Michael Burczynski, NIH-TrialBench Consortium, Zhiyong Lu
arXiv:2608.30022v1 Announce Type: new
Abstract: Introduction: NICE guidelines provide evidence-based recommendations for clinical care but remain largely in unstructured natural language. Existing ap...
By Ashvin Gupta, Denys Prociuk, Alessandra Russo, Brendan C. Delaney
The paper introduces VERGE, a verification-enhanced refinement workflow that extracts six red‑flag symptoms and family‑history risk status for early‑onset colorectal cancer from free‑text clinical notes. VERGE uses retrieval‑augmented generation followed by a bounded verification‑refinement cycle that checks textual grounding and clinical validity, correcting claims until resolved or escalating to human review. In evaluation on 4,033 clinician‑labeled note‑finding pairs, VERGE improved precision from 0.764 to 0.849 and MCC from 0.681 to 0.730 compared to a single‑agent baseline, while requiring human review for only 1.5 % of claims.
By Nikkie Hooman, Monarch Nigam, Amy E. Hughes, Rasmi G. Nair, Mehak Gupta
The paper introduces DeToxR, a reinforcement‑learning‑enhanced large language model designed to support decision making in acute toxicology cases. It fuses unstructured narratives from paramedics and patients with structured vital‑sign data to predict co‑ingested substances across 14 classes. In preliminary validation, DeToxR outperforms baseline models, achieving higher micro‑F1 and recall scores for poison identification.
By Nico Oberl\"ander, David Bani-Harouni, Tobias Zellner, Nassir Navab, Florian Eyer, Matthias Keicher