arXiv AI

T2D-Bench: Evidence-Gated Evaluation of LLM Outputs for Type 2 Diabetes Using a Multi-Layer Clinical-Lifestyle Knowledge Graph

arXiv:2606. 24145v1 Announce Type: new Abstract: Large language models (LLMs) can produce clinically fluent recommendations for type 2 diabetes while failing to satisfy guideline constraints or explicitly justify lifestyle-related glycemic claims.

arXiv Computation and Language
Sep 1

DIASENTINEL: An Auditable Multi-Agent System for Guideline-Grounded Diabetes Risk Screening

arXiv:2608.31128v1 Announce Type: new Abstract: Large language models (LLMs) offer promising clinical decision support but remain vulnerable to hallucinated facts, unsupported recommendations, and ci...

By Yung Wei Shueh, Zhi-Jie Chen, Chia-Hsuan Hsu, Hsin-Ling Hsu, Donghua Zhang, Chenwei Wu, Jun-En Ding, Tongze Zhang, Shihao Yang, Pengfei Hu, Fang-Ming Hung, Feng Liu
arXiv AI
Sep 10

Building evidence-based knowledge bases from full-text literature for disease-specific biomedical reasoning

EvidenceNet is a disease‑specific dataset that transforms full‑text biomedical literature into structured evidence records and graph representations, preserving study design, provenance, and quantitative support. Using an LLM‑assisted pipeline, it extracts experimentally grounded findings, normalizes entities, scores evidence quality, and links related records via typed semantic relations. The released subsets—EvidenceNet‑HCC and EvidenceNet‑CRC—contain thousands of evidence records and richly connected graphs, with high extraction and relation‑type accuracy, enabling retrieval‑augmented question answering and graph‑based tasks such as link prediction and target prioritization.

By Chang Zong, Jinyu Chen, Sicheng Lv, Si-tu Xue, Huilin Zheng, Jian Wan, Lei Zhang
arXiv Computation and Language
Sep 23

Quantitative Evidence Mining for Plausibility-Aware Biomedical AI: A Narrative Review and Conceptual Framework

The article proposes a framework called quantitative evidence mining to transform biomedical findings into structured, context-rich evidence units. It outlines core elements such as claim, measured entity, value, comparator, population, conditions, temporal context, uncertainty, provenance, validation, and expert review. The authors present an eight-stage reference architecture and emphasize that plausibility should remain multidimensional rather than collapsed into a single truth label, linking extraction to evidence synthesis for applications like clinical trials, biomarker research, and knowledge-graph construction.

By Negin Sadat Babaiha, Stefan Geissler, Marie-Christine Simon, Martin Hofmann-Apitius, Marc Jacobs
Hugging Face Trending Papers
Sep 8

It's All in the Way You Say It: The Role of Information Representation in LLM-Based Glycemic-Event Prediction

The paper examines how the representation of physiological data affects the performance of large language models (LLMs) in predicting post‑meal blood glucose events for people with type 1 diabetes. Using the OhioT1DM dataset, the authors compare zero‑shot and few‑shot prompt‑based LLMs across 30, 60, and 90‑minute horizons, varying the textual encoding of glucose readings, derived descriptors, and contextual variables such as insulin, meals, carbs, and activity. Results show that while conventional supervised models excel at hyperglycemia prediction, certain prompt‑based LLM configurations outperform them for hypoglycemia, and that the way data is presented to the model is a key determinant of success, with added context not consistently improving outcomes.

arXiv AI
2d ago

OpenMTB-Audit: Exposing Over-Refusal and Clinical Expert Perspectives in LLM-Based Molecular Tumor Board Safety Evaluation

OpenMTB‑Audit is an open‑source benchmark that tests large language models on 500 synthetic non‑small cell lung cancer cases, covering five adversarial error categories and four safety labels: Supported, Partially Supported, Unsupported, and Insufficient Information. The study found that all eight tested LLMs over‑refused Partially Supported recommendations, collapsing labels to achieve high safety scores. A deterministic seven‑module framework, MTB‑AuditAgent, was introduced to reduce over‑refusal to 6.7% and reach 91.2% accuracy, while an oncologist annotation study highlighted disagreement around the boundary between information sufficiency and treatment optimization.

By Negin Ashrafi, Jia Luo, Stacey M. Frumm, Roxana Daneshjou