arXiv AI By Manan Roy Choudhury, Suparno Roy Chowdhury, Swastik Sahoo, Muhammad Ali Khan, Kaneez Zahra Rubab Khakwani, Mohamad Bassam Sonbol, Irbaz Bin Riaz, Vivek Gupta

Answering clinicians' questions over trial evidence tables with verifiable, feedback-driven language models

Read the original on arXiv AI →

FD‑SCoPE is a language‑model framework that answers clinicians’ questions about systematic review evidence tables, exposing the underlying query, selected trials, and derivation rule for each answer. It handles both directly recorded attributes and derived attributes, achieving high accuracy on an oncology evidence table of 159 immune‑checkpoint inhibitor trials. After incorporating expert corrections, its performance on unseen questions improved from 77.9% to 84.9% F1.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 14

Information-seeking failures of large language models in agentic clinical reasoning

arXiv:2607. 10275v1 Announce Type: new Abstract: Large language models achieve high scores on medical knowledge assessments, yet clinical reasoning requires actively deciding what to investigate under uncertainty.

By Krischan Braitsch, Laura K. Schmalbrock, Theresa Weltermann, Andrew F. Berdel, Isabella Miller, Kai Tran, Michael Heider, Sabrina Kraus, Florian Bassermann, Jacqueline Lammert, Sebastian Ziegelmayer, Marcus Makowski, Lisa C. Adams, Keno K. Bressem
arXiv AI
Sep 12

Can LLMs Follow Medical Expert Logic? A Benchmark for Hierarchical Logical Consistency in Risk-of-Bias Assessment

LogiMed‑RoB is a new benchmark that tests large language models (LLMs) on hierarchical logical consistency in medical risk‑of‑bias assessments, using 860 randomized controlled trials and 14,820 queries based on Cochrane Risk of Bias 2.0 expert logic. The benchmark evaluates models across four dimensions—Atomic Consistency, Domain Consistency, Aggregation Consistency, and Evidential Faithfulness—revealing a catastrophic error‑compounding effect where high atomic accuracy does not translate to end‑to‑end consistency. Experiments on ten state‑of‑the‑art LLMs show that even top models can fail to deduce correct outcomes in a significant portion of cases, highlighting a gap between evidence retrieval and reasoning. whyItMatters":"The study shows that high outcome accuracy can mask critical reasoning flaws, emphasizing the need for white‑box logical verification before deploying LLMs in clinical settings."

By Jiayu Huang, Zichen Tang, Qianhui Ling, Zemin Kuang, Haihong E
Hugging Face Trending Papers
Sep 10

Can LLMs Follow Medical Expert Logic? A Benchmark for Hierarchical Logical Consistency in Risk-of-Bias Assessment

The paper introduces LogiMed‑RoB, a benchmark that tests large language models (LLMs) on hierarchical logical consistency in medical risk‑of‑bias assessments, using 860 randomized controlled trials and 14,820 queries based on Cochrane RoB 2.0 expert logic. It evaluates models across four dimensions—Atomic Consistency, Domain Consistency, Aggregation Consistency, and Evidential Faithfulness—revealing a severe Error Compounding Effect where high atomic accuracy does not translate to end‑to‑end consistency. Experiments on ten state‑of‑the‑art LLMs show that even top performers can collapse to 45.13% overall consistency, with some models nearly failing entirely, and that many models struggle to deduce correct outcomes from retrieved evidence.

arXiv AI
Sep 3

FDARxBench: Benchmarking Regulatory and Clinical Reasoning on FDA Generic Drug Assessment

FDARxBench is an expert‑curated benchmark designed to evaluate document‑grounded question answering on FDA drug label documents, focusing on generic drug assessment. It features a multi‑stage pipeline that generates high‑quality QA examples covering factual, multi‑hop, and refusal tasks, and includes protocols for both open‑book and closed‑book reasoning. Experiments with various language models show significant gaps in factual grounding, long‑context retrieval, and safe refusal behavior, highlighting the challenge of regulatory‑grade label comprehension.

By Betty Xiong, Jillian Fisher, Benjamin Newman, Meng Hu, Shivangi Gupta, Yejin Choi, Lanyan Fang, Russ B Altman