arXiv AI

What Can Artificial Intelligence Learn from Medicine? Generative Analogies and Reliable Machine Learning Systems

The paper examines how machine learning (ML) in medicine parallels clinical translation, framing this relationship as a generative analogy. It identifies epistemic and methodological warrants from clinical translation that are often invoked in ML discussions and analyzes how these warrants apply analogically to ML. The authors propose a new form of ML reliabilism, interpreting clinical translation warrants in reliabilist terms and showing how this perspective can inform ML systems distinct from existing reliabilist accounts in philosophy of AI.

arXiv AI
Aug 13

Teaching agentic AI to learn expert reasoning for rare disease diagnosis

arXiv:2606. 16149v3 Announce Type: replace Abstract: Rare disease diagnosis depends on expert reasoning that is scarce and difficult to transfer; off-the-shelf large language models (LLMs) rank the correct disease first in only 35.

By Minh-Ha Nguyen, Erica Gray, Bryce A. Schuler, Kevin W. Byram, Chih-Ting Yang, Fan Ma, Hua Xu, Wu-Chen Su, Chao Yan, Wei-Qi Wei, Adam Wright, Lisa Bastarache, Josh Peterson, Lingyao Li, Siyuan Ma, Undiagnosed Diseases Network, Rizwan Hamid, Thomas A. Cassini, Cathy Shyr
arXiv AI
Sep 10

Teaching agentic AI to generalize expert diagnostic reasoning in rare diseases

arXiv:2606.16149v5 Announce Type: replace Abstract: Rare disease diagnosis depends on expert reasoning that is scarce and difficult to transfer. Large language models rank the correct disease first i...

By Minh-Ha Nguyen, Erica Gray, Bryce A. Schuler, Kevin W. Byram, Chih-Ting Yang, Fan Ma, Hua Xu, Wu-Chen Su, Chao Yan, Wei-Qi Wei, Adam Wright, Lisa Bastarache, Josh F. Peterson, Lingyao Li, Siyuan Ma, Undiagnosed Diseases Network, Rizwan Hamid, Thomas A. Cassini, Cathy Shyr
arXiv Machine Learning
Sep 14

Language Is an Insufficient Substrate for Quantitative Reasoning, and Consequential Domains Need Large Quantitative Models

The article argues that large language models (LLMs) are inadequate for consequential quantitative tasks such as pricing, risk assessment, and medical triage because language is a lossy representation of quantitative data that cannot be reversed. It formalizes this limitation as a property of the training representation rather than model capacity and identifies three essential properties—reproducibility, traceable lineage to source records, and calibrated uncertainty—that language substrates cannot provide. The authors propose a new class of models, Large Quantitative Models (LQMs), designed to meet these requirements.

By Reuben Vandeventer, David Imrem, David J. Wild
arXiv AI
Aug 18

Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning

arXiv:2608. 14804v1 Announce Type: new Abstract: Large language models (LLMs) have become the dominant interface of clinical artificial intelligence, yet the interface they expose (text in, text out, one context window at a time) maintains no explicit, persistent, governed representation of what is currently true about a patient.

By Augusto Bernardo Pissarra, Victor Lorena de Farias Souza
arXiv AI
Jun 2

Algorithmic Authority and the Clinical Standard of Care

arXiv:2606. 00044v1 Announce Type: cross Abstract: The integration of artificial intelligence into clinical medicine creates a fundamental tension between algorithmic probabilistic reasoning and the experiential intuition of expert physicians; applying Lawrence Lessig's \enquote{Code is Law} framework, I argue that the architecture of clinical AI systems already functions as de facto medical regulation, reshaping liability and the standard of care.

By Aizierjiang Aiersilan
arXiv AI
Sep 4

Medical Reasoning in the Era of LLMs: A Systematic Review of Enhancement Techniques and Applications

The paper reviews how Large Language Models (LLMs) are being adapted for medical reasoning, moving beyond single-step answers to systems that can systematically, transparently, and verifiably reason. It introduces a taxonomy of enhancement techniques, split into training-time methods such as supervised fine‑tuning and reinforcement learning, and test-time methods like prompt engineering and multi‑agent systems. The review examines their application across text, image, and code modalities in key clinical areas—diagnosis, education, and treatment planning—and tracks the shift in evaluation benchmarks from simple accuracy to more nuanced assessments of reasoning quality and visual interpretability.

By Zizhan Ma, Wenxuan Wang, Meidan Ding, Shiyi Zheng, Shengyuan Liu, Jie Liu, Jiaming Ji, Linlin Shen, Yixuan Yuan, Wenting Chen