Can Jev Judge Radiology Reports? Evaluating a System One Model for Clinical Factuality
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2609.01470v1 Announce Type: new Abstract: As AI systems are increasingly used to draft radiology reports, reliably evaluating their clinical quality remains a critical challenge. Large language...
arXiv:2606. 17062v1 Announce Type: cross Abstract: Radiology report evaluation must distinguish clinical compatibility from surface similarity, because negation, laterality, or normal-abnormal polarity can reverse a finding.
arXiv:2606. 08769v1 Announce Type: cross Abstract: Automatic evaluation is critical for high-stakes text generation, where errors often involve omitted findings, hallucinated content, polarity reversals, location changes, uncertainty mismatches, and temporal-comparison errors rather than low surface similarity alone.
The study examines how differences in radiologists’ reporting styles—such as terminology, shorthand, formatting, and detail—affect the evaluation of AI-generated chest X‑ray reports. By quantifying the sensitivity of common metrics to these variations, the authors show that changes in reference reports can shift model rankings. They introduce a taxonomy of reporting variations and a rewriting method, ReRef, that preserves clinical meaning while altering style, and release a validated dataset of paired reference reports to aid future research.
arXiv:2608.31016v1 Announce Type: cross Abstract: Ambient AI scribes draft clinical notes, and published audits find their dominant error is omission: information the encounter established that the n...
arXiv:2604. 14892v3 Announce Type: replace-cross Abstract: Evaluating medical AI systems using expert clinician panels is costly and slow, motivating the use of large language models (LLMs) as alternative adjudicators.