arXiv:2608. 05188v1 Announce Type: cross Abstract: Despite ever-increasing sophistication in language model (LM) pre- and post-training pipelines, many important failures persist: models overcondition on user framing ("sycophancy"), exhibit incomplete logical generalization, and produce confident but incorrect responses.
By Itamar Pres, Belinda Z. Li, Laura Ruis, Zifan Carl Guo, Keya Hu, Mehul Damani, Isha Puri, Ekdeep Singh Lubana, Jacob Andreas
arXiv:2507. 02778v3 Announce Type: replace-cross Abstract: Although large language models (LLMs) have transformed AI, they still make errors and follow unproductive reasoning paths.
By Ken Tsui
arXiv:2607. 08065v1 Announce Type: new Abstract: LLM-as-judge (Zheng et al.
By Kaihua Ding
LogiMed‑RoB is a new benchmark that tests large language models (LLMs) on hierarchical logical consistency in medical risk‑of‑bias assessments, using 860 randomized controlled trials and 14,820 queries based on Cochrane Risk of Bias 2.0 expert logic. The benchmark evaluates models across four dimensions—Atomic Consistency, Domain Consistency, Aggregation Consistency, and Evidential Faithfulness—revealing a catastrophic error‑compounding effect where high atomic accuracy does not translate to end‑to‑end consistency. Experiments on ten state‑of‑the‑art LLMs show that even top models can fail to deduce correct outcomes in a significant portion of cases, highlighting a gap between evidence retrieval and reasoning.
whyItMatters":"The study shows that high outcome accuracy can mask critical reasoning flaws, emphasizing the need for white‑box logical verification before deploying LLMs in clinical settings."
By Jiayu Huang, Zichen Tang, Qianhui Ling, Zemin Kuang, Haihong E
The paper introduces LogiMed‑RoB, a benchmark that tests large language models (LLMs) on hierarchical logical consistency in medical risk‑of‑bias assessments, using 860 randomized controlled trials and 14,820 queries based on Cochrane RoB 2.0 expert logic. It evaluates models across four dimensions—Atomic Consistency, Domain Consistency, Aggregation Consistency, and Evidential Faithfulness—revealing a severe Error Compounding Effect where high atomic accuracy does not translate to end‑to‑end consistency. Experiments on ten state‑of‑the‑art LLMs show that even top performers can collapse to 45.13% overall consistency, with some models nearly failing entirely, and that many models struggle to deduce correct outcomes from retrieved evidence.
arXiv:2608. 09080v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved strong performance in medical question answering and clinical reasoning tasks.
By Maryam Tahermazandarani, Adnan Mahmood, Fahmida Islam, Quan Z. Sheng
arXiv:2607. 24371v1 Announce Type: cross Abstract: Healthcare interoperability requires AI systems to produce structured outputs conforming to standardized schemas including ICD-10 for diagnostic coding, CPT for procedure billing, and HL7 FHIR for data exchange.
By Jianru Shen
arXiv:2607. 28677v1 Announce Type: new Abstract: LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning.
By Shayndhan Sivanathan, Shravan Nageswaran, Mehdi Zadem, Ryaan Sultan, Nicolas von Mallinckrodt, Max Solovyev, Alexey Matyushkin, Sumon Sadhu, Gabriele C DeLuca, Sanjeeva Jeyaretna, James Hillis, Manoj Ramachandran, Prakash Jayakumar
The study evaluates large language models (LLMs) on sequential emergency department triage, where acuity labels are predicted from progressively longer nurse‑patient conversations. Six LLMs were tested at five checkpoints on simulated and physician‑authored dialogues, showing a decline from moderate‑to‑substantial agreement on full records to only fair‑to‑moderate agreement at each checkpoint. The models consistently anchor on chief complaint exchanges and fail to integrate later evidence, yielding low agreement with clinicians (QWK 0.295 vs. 0.887‑0.929) and concentrating predictions on ESI‑2 and ESI‑3.
whyItMatters":"The findings reveal that LLMs, despite strong offline performance, cannot reliably handle the sequential nature of real‑time triage, highlighting a critical gap for safe deployment in emergency settings."
By Dipankar Srirag, Haokai Zhao, Ashutosh Kumar, Eleanor Hopper, Michael Dalton, Quoc Dung Nguyen, Aditya Joshi, Salil S. Kanhere, Padmanesan Narasimhan
The paper introduces a decision‑theoretic framework that elicits both probability judgments and decisions from large language models (LLMs) to test whether their reported beliefs are consistent with their actions. It shows that this framework yields empirically testable conditions without assuming a specific utility function. In clinical diagnosis simulations, the authors find that while LLMs’ reported beliefs are not perfect reflections of the information in their decisions, the discrepancies are small for the strongest models.
By Khurram Yamin, Jingjing Tang, Santiago Cortes-Gomez, Amit Sharma, Eric Horvitz, Bryan Wilder
Consistency training encourages a model to produce similar outputs across related inputs or sampling procedures. Such methods are simple, scalable, and largely label-free, but their effects on model alignment remain poorly understood.
The paper introduces an LLM-as-a-judge method to quantify perturbation strength for assessing self-consistency in LLM-generated explanations. It evaluates explanations from various LLMs under controlled perturbation strengths, comparing input and chain-of-thought (CoT) perturbations. Results show the proposed measure outperforms embedding- and probability-based approaches, and that input perturbations impact LLMs more than CoT perturbations, implying fair self-consistency judgments require consistent perturbation types.
By Phuong Q. Le, Kemal Kurniawan, Jey Han Lau