Hugging Face Trending Papers

Can LLMs Follow Medical Expert Logic? A Benchmark for Hierarchical Logical Consistency in Risk-of-Bias Assessment

Read the original on Hugging Face Trending Papers →

The paper introduces LogiMed‑RoB, a benchmark that tests large language models (LLMs) on hierarchical logical consistency in medical risk‑of‑bias assessments, using 860 randomized controlled trials and 14,820 queries based on Cochrane RoB 2.0 expert logic. It evaluates models across four dimensions—Atomic Consistency, Domain Consistency, Aggregation Consistency, and Evidential Faithfulness—revealing a severe Error Compounding Effect where high atomic accuracy does not translate to end‑to‑end consistency. Experiments on ten state‑of‑the‑art LLMs show that even top performers can collapse to 45.13% overall consistency, with some models nearly failing entirely, and that many models struggle to deduce correct outcomes from retrieved evidence.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
6d ago

Can LLMs Follow Medical Expert Logic? A Benchmark for Hierarchical Logical Consistency in Risk-of-Bias Assessment

LogiMed‑RoB is a new benchmark that tests large language models (LLMs) on hierarchical logical consistency in medical risk‑of‑bias assessments, using 860 randomized controlled trials and 14,820 queries based on Cochrane Risk of Bias 2.0 expert logic. The benchmark evaluates models across four dimensions—Atomic Consistency, Domain Consistency, Aggregation Consistency, and Evidential Faithfulness—revealing a catastrophic error‑compounding effect where high atomic accuracy does not translate to end‑to‑end consistency. Experiments on ten state‑of‑the‑art LLMs show that even top models can fail to deduce correct outcomes in a significant portion of cases, highlighting a gap between evidence retrieval and reasoning. whyItMatters":"The study shows that high outcome accuracy can mask critical reasoning flaws, emphasizing the need for white‑box logical verification before deploying LLMs in clinical settings."

By Jiayu Huang, Zichen Tang, Qianhui Ling, Zemin Kuang, Haihong E
arXiv Computer Vision
Aug 24

Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective Refinement

arXiv:2608.10725v2 Announce Type: replace Abstract: Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing mo...

By Uma Ranjan, Kunal Tilaganji, Aditya Koul, Anurag Mahipal, Dashpreet Singh, Hriday Rana, Manan Jain, Sidharth Gupta, Ajo Babu George, Vineeth Balasubramanian, Nagarajan Natarajan, Amit Sharma
arXiv AI
Jul 14

Information-seeking failures of large language models in agentic clinical reasoning

arXiv:2607. 10275v1 Announce Type: new Abstract: Large language models achieve high scores on medical knowledge assessments, yet clinical reasoning requires actively deciding what to investigate under uncertainty.

By Krischan Braitsch, Laura K. Schmalbrock, Theresa Weltermann, Andrew F. Berdel, Isabella Miller, Kai Tran, Michael Heider, Sabrina Kraus, Florian Bassermann, Jacqueline Lammert, Sebastian Ziegelmayer, Marcus Makowski, Lisa C. Adams, Keno K. Bressem
arXiv AI
Jun 15

Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning?

arXiv:2604. 14892v3 Announce Type: replace-cross Abstract: Evaluating medical AI systems using expert clinician panels is costly and slow, motivating the use of large language models (LLMs) as alternative adjudicators.

By Amy Rouillard, Sitwala Mundia, Linda Camara, Ziyaad Dangor, Michael Cameron Gramanie, Ismail Kalla, Shabir A. Madhi, Kajal Morar, Marlvin T. Ncube, Haroon Saloojee, Bruce A. Bassett
arXiv AI
Sep 10

Teaching agentic AI to generalize expert diagnostic reasoning in rare diseases

arXiv:2606.16149v5 Announce Type: replace Abstract: Rare disease diagnosis depends on expert reasoning that is scarce and difficult to transfer. Large language models rank the correct disease first i...

By Minh-Ha Nguyen, Erica Gray, Bryce A. Schuler, Kevin W. Byram, Chih-Ting Yang, Fan Ma, Hua Xu, Wu-Chen Su, Chao Yan, Wei-Qi Wei, Adam Wright, Lisa Bastarache, Josh F. Peterson, Lingyao Li, Siyuan Ma, Undiagnosed Diseases Network, Rizwan Hamid, Thomas A. Cassini, Cathy Shyr