arXiv AI By Jiayu Huang, Zichen Tang, Qianhui Ling, Zemin Kuang, Haihong E

Can LLMs Follow Medical Expert Logic? A Benchmark for Hierarchical Logical Consistency in Risk-of-Bias Assessment

Read the original on arXiv AI →

LogiMed‑RoB is a new benchmark that tests large language models (LLMs) on hierarchical logical consistency in medical risk‑of‑bias assessments, using 860 randomized controlled trials and 14,820 queries based on Cochrane Risk of Bias 2.0 expert logic. The benchmark evaluates models across four dimensions—Atomic Consistency, Domain Consistency, Aggregation Consistency, and Evidential Faithfulness—revealing a catastrophic error‑compounding effect where high atomic accuracy does not translate to end‑to‑end consistency. Experiments on ten state‑of‑the‑art LLMs show that even top models can fail to deduce correct outcomes in a significant portion of cases, highlighting a gap between evidence retrieval and reasoning. whyItMatters":"The study shows that high outcome accuracy can mask critical reasoning flaws, emphasizing the need for white‑box logical verification before deploying LLMs in clinical settings."

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Sep 10

Can LLMs Follow Medical Expert Logic? A Benchmark for Hierarchical Logical Consistency in Risk-of-Bias Assessment

The paper introduces LogiMed‑RoB, a benchmark that tests large language models (LLMs) on hierarchical logical consistency in medical risk‑of‑bias assessments, using 860 randomized controlled trials and 14,820 queries based on Cochrane RoB 2.0 expert logic. It evaluates models across four dimensions—Atomic Consistency, Domain Consistency, Aggregation Consistency, and Evidential Faithfulness—revealing a severe Error Compounding Effect where high atomic accuracy does not translate to end‑to‑end consistency. Experiments on ten state‑of‑the‑art LLMs show that even top performers can collapse to 45.13% overall consistency, with some models nearly failing entirely, and that many models struggle to deduce correct outcomes from retrieved evidence.

arXiv Computer Vision
Aug 24

Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective Refinement

arXiv:2608.10725v2 Announce Type: replace Abstract: Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing mo...

By Uma Ranjan, Kunal Tilaganji, Aditya Koul, Anurag Mahipal, Dashpreet Singh, Hriday Rana, Manan Jain, Sidharth Gupta, Ajo Babu George, Vineeth Balasubramanian, Nagarajan Natarajan, Amit Sharma
arXiv AI
Jul 14

Information-seeking failures of large language models in agentic clinical reasoning

arXiv:2607. 10275v1 Announce Type: new Abstract: Large language models achieve high scores on medical knowledge assessments, yet clinical reasoning requires actively deciding what to investigate under uncertainty.

By Krischan Braitsch, Laura K. Schmalbrock, Theresa Weltermann, Andrew F. Berdel, Isabella Miller, Kai Tran, Michael Heider, Sabrina Kraus, Florian Bassermann, Jacqueline Lammert, Sebastian Ziegelmayer, Marcus Makowski, Lisa C. Adams, Keno K. Bressem
arXiv AI
Jun 15

Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning?

arXiv:2604. 14892v3 Announce Type: replace-cross Abstract: Evaluating medical AI systems using expert clinician panels is costly and slow, motivating the use of large language models (LLMs) as alternative adjudicators.

By Amy Rouillard, Sitwala Mundia, Linda Camara, Ziyaad Dangor, Michael Cameron Gramanie, Ismail Kalla, Shabir A. Madhi, Kajal Morar, Marlvin T. Ncube, Haroon Saloojee, Bruce A. Bassett
arXiv AI
Jul 10

A safety-oriented hypothetico-deductive framework for AI-assisted differential diagnosis

arXiv:2607. 08038v1 Announce Type: new Abstract: Diagnostic error is a major threat to patient safety, yet current large language model (LLM) systems often treat diagnosis as a one-shot prediction task, lacking safeguards against missed high-risk alternatives or rigorous verification of their reasoning.

By Fan Ma, Mauro Giuffr\`e, Donald Wright, Kent McCann, Mark Iscoe, Lingfei Qian, Mingyang Jiang, Chi Wing Ng, Na Hong, Huan He, Cathy Shyr, Qingyu Chen, Lee Schwamm, Lucila Ohno-Machado, Hua Xu