arXiv Computation and Language By Zikai Zhou, Yufei Jin, Yilin Xu, Yu-Chiang Wang, Chieh-Ju Chao, Monica S. Lam

Accountable AI with Grounded, Faithful, Consistent, Actionable Rationales: A Case Study in Clinical Trial Matching with VERDICT

Read the original on arXiv Computation and Language →

The paper introduces VERDICT, an LLM-based agent that converts clinical trial matching tasks into SMT problems to ensure consistent policy application and accountable decisions. VERDICT outperforms other LLM-only and neurosymbolic baselines on accuracy, achieves perfect policy consistency, and generates clinician-preferred rationales grounded in explicit assumptions and pivotal conditions. It also demonstrates improved counterfactual self‑faithfulness, meaning changes in pivotal conditions appropriately alter decisions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jun 15

Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning?

arXiv:2604. 14892v3 Announce Type: replace-cross Abstract: Evaluating medical AI systems using expert clinician panels is costly and slow, motivating the use of large language models (LLMs) as alternative adjudicators.

By Amy Rouillard, Sitwala Mundia, Linda Camara, Ziyaad Dangor, Michael Cameron Gramanie, Ismail Kalla, Shabir A. Madhi, Kajal Morar, Marlvin T. Ncube, Haroon Saloojee, Bruce A. Bassett
arXiv AI
Aug 24

No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators

The paper introduces a framework for evaluating AI systems that not only checks final labels but also tracks the reasoning behind them through three core sources—grounds, norms, and authority—forming an eight-cell counterfactual judgment cube. It defines minimal source replacement sets, called judgment receipts, to explain changes in verdicts and provides certification cost bounds for black-box evaluators. The authors present ReasonBench, a benchmark with 19,520 cases, and demonstrate that while high standard accuracy can mask robustness issues, receipt accuracy reveals significant gaps in reasoning consistency across different models.

By Ye Chen, Weining Zhang
arXiv AI
Sep 7

Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?

The paper proposes a new method for evaluating AI accountability by analyzing the structural quality of a model’s defense for its decisions, using a four‑phase dialectical protocol based on Walton’s argumentation schemes and Govier’s criteria. Applied to nine large language models and 200 ambiguous moral-choice items, the study finds that models generally defend their reasoning well above the rubric minimum, though failures cluster on grounds and sufficiency and correlate with epistemic hedging. The protocol also reveals that models often present different argument schemes in justification than in reasoning, detects indefensible defenses, and highlights challenges in assessing retraction in AI alignment.

By Daan R. Henselmans, Derck W. E. Prinzhorn, Arno Libert
arXiv AI
Sep 12

Can LLMs Follow Medical Expert Logic? A Benchmark for Hierarchical Logical Consistency in Risk-of-Bias Assessment

LogiMed‑RoB is a new benchmark that tests large language models (LLMs) on hierarchical logical consistency in medical risk‑of‑bias assessments, using 860 randomized controlled trials and 14,820 queries based on Cochrane Risk of Bias 2.0 expert logic. The benchmark evaluates models across four dimensions—Atomic Consistency, Domain Consistency, Aggregation Consistency, and Evidential Faithfulness—revealing a catastrophic error‑compounding effect where high atomic accuracy does not translate to end‑to‑end consistency. Experiments on ten state‑of‑the‑art LLMs show that even top models can fail to deduce correct outcomes in a significant portion of cases, highlighting a gap between evidence retrieval and reasoning. whyItMatters":"The study shows that high outcome accuracy can mask critical reasoning flaws, emphasizing the need for white‑box logical verification before deploying LLMs in clinical settings."

By Jiayu Huang, Zichen Tang, Qianhui Ling, Zemin Kuang, Haihong E