Hugging Face Trending Papers

Beyond Majority Vote: Multi-Perspective Adjudication for Medical Hallucination Detection

The paper investigates how to better detect factual errors, or hallucinations, in long-form medical chatbot responses. It introduces a multi‑perspective annotation workflow that combines first‑pass labeling, a large language model acting as a judge (LaJ) to surface candidate errors, and two adjudication steps—expert medical review and evidence‑based fact‑checking. The study finds that single‑pass benchmarks miss many errors, that LaJ alone is insufficient, and that adjudicators disagree, indicating that multi‑pass adjudication improves coverage but still depends on human judgment and evidence.

arXiv Computation and Language
Sep 4

Beyond Majority Vote: Multi-Perspective Adjudication for Medical Hallucination Detection

The paper presents a multi‑perspective annotation framework for detecting medical hallucinations in chatbot responses. It combines first‑pass annotators, a large language model acting as a judge (LaJ) for candidate discovery, and two adjudication stages—medical‑expert review and evidence‑based fact‑checking. The study finds that single‑pass labeling undercounts errors, while multi‑pass adjudication improves coverage but still depends on expert judgment and evidence.

By Joe Cecil, Marjorie Freedman
arXiv AI
1d ago

Scaling Clinical Judgment to Evaluate Medical AI

arXiv:2609.12822v2 Announce Type: replace Abstract: Blinded physician evaluation has been considered by many to be the gold standard for assessing clinical reasoning in large language models (LLMs)....

By Thomas A. Buckley, Zahir Kanjee, Peter G. Brodeur, Byron Crowe, Anthony M. Pettinato, Aashna P. Shah, Adrian D. Haimovich, Liam G. McCoy, Daniel Restrepo, Jason A. Freed, Ethan Goh, Jonathan H. Chen, Laura Zwaan, Katherine E. Goodman, Daniel J. Morgan, Raja-Elie E. Abdulnour, Adam Rodman, Arjun K. Manrai
arXiv Computation and Language
5d ago

Where Does Retrieval-Based Open-Ended Evaluation Fail? Automatic Taxonomy Induction from Long-Form Medical Answer Factuality Verification

The paper investigates why retrieval‑based open‑ended evaluation fails in medical fact verification. By creating two detailed taxonomies—one for retrieval‑stage errors across five quality dimensions and another for verifier‑reasoning errors across six steps—the authors automatically label evidence quality and reasoning errors using an LLM‑as‑Judge pipeline. Their large‑scale stress tests across multiple retrieval methods and verifier models show that increasing model size, reasoning effort, source breadth, or medical fine‑tuning does not eliminate these failure modes, indicating fundamental limits of the retrieve‑then‑verify paradigm in open‑ended medical contexts.

By Heyuan Huang, Jirui Dai, Alexandra DeLucia, Sonal Joshi, Mahsa Yarmohammadi, Jie Gao, Bernal Jim\'enez Guti\'errez, Mark Dredze
arXiv AI
Aug 28

MedFabric: Gold Evidence Hides the Difficulty of Word-Level Medical Fabrication Detection

MedFabric is a new benchmark for detecting word‑level medical fabrications, comprising 646 fabricated statements each paired with a ground‑truth passage that shares the same LLM authorship and nearly identical wording. The study shows that current detectors perform poorly—expert clinicians achieve only 53.3% macro‑F1 and no detector family surpasses 60% without gold evidence—highlighting that detection hinges on evidence correctness rather than subtlety of fabrication. The authors demonstrate that a retrieval‑confidence gate can substantially improve performance, raising macro‑F1 from 61% to 74%.

By Tung Sum Thomas Kwok, Qian Qian, Xiaofeng Lin, Dongxu Zhang, Jun Han, Zhichao Yang, Davin Hill, Tamer Soliman, Sanjit Singh Batra, Robert Tillman, Guang Cheng
arXiv AI
Jun 15

Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning?

arXiv:2604. 14892v3 Announce Type: replace-cross Abstract: Evaluating medical AI systems using expert clinician panels is costly and slow, motivating the use of large language models (LLMs) as alternative adjudicators.

By Amy Rouillard, Sitwala Mundia, Linda Camara, Ziyaad Dangor, Michael Cameron Gramanie, Ismail Kalla, Shabir A. Madhi, Kajal Morar, Marlvin T. Ncube, Haroon Saloojee, Bruce A. Bassett
arXiv AI
Jul 29

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

arXiv:2607. 25485v1 Announce Type: new Abstract: Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf.

By Korosh Vatanparvar, Ashutosh Joshi, Maria Xenochristou, Mohammad Abuzar Hashemi, Prasad Kasu, Deepak Bansal, Daniel Lopez-Martinez, Anchal Nema, Ramya Ganesan, Will Kimbrough, Alex Woody, Yadunandana Rao, Dilek Hakkani-Tur, Wilko Schulz-Mahlendorf
arXiv Machine Learning
Aug 24

Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment

The paper introduces an LLM-as-a-Judge framework for evaluating the outputs of an agentic drug discovery assistant, ChatInvent, deployed at AstraZeneca. It defines four quality dimensions—Completeness, Relevancy, Structural Clarity, and Scope Adherence—alongside deterministic Tool Call Correctness checks, and validates the judge against five expert annotators. After optimizing the best-performing judge with few-shot demonstrations, alignment with human majority votes improves from 0.80 to 0.86, and the framework reveals that informal question phrasing does not degrade output quality.

By Emma Granqvist, Roc\'io Mercado, Samuel Genheden
arXiv Computation and Language
Sep 14

Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models

The paper introduces Tasks over Application Manuals (TAM), a benchmark designed to test long‑horizon procedural reasoning in large language models. TAM uses real‑world tasks from ICD‑10‑CM clinical coding and U.S. federal sentencing, requiring models to follow extensive, rule‑based manuals and perform interdependent steps to produce exact answers. Experiments with GPT‑5 and various prompting strategies show very low exact‑match accuracy—1% for coding and 15.5% for sentencing—highlighting a gap between current benchmarks and the ability to reliably follow complex procedures.

By Utkarsh Soni, Syed Shariyar Murtaza, Yifan Nie, Sachin Chandrasekhar, Eugene Wen