Corpus Characterization and Inverse Constitutional Fine-Tuning for Style-Aware Radiology Reports
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
ASTAR is an LLM-based framework that automatically generates standardized radiology reporting templates from large-scale clinical free-text corpora, eliminating the manual, expert-driven template construction process. In experiments on 4,215 fetal brain MRI reports from multiple centers, ASTAR‑induced templates outperformed two expert‑curated templates in template coverage, information fidelity, diagnostic fidelity, and expert‑rated usability. The approach reduces template development time from weeks of committee deliberation to hours of automated processing.
arXiv:2607. 05880v1 Announce Type: cross Abstract: Imaging demand is growing faster than the radiology workforce can expand, and reporting backlogs cannot be resolved through training and recruitment alone.
The study examines how differences in radiologists’ reporting styles—such as terminology, shorthand, formatting, and detail—affect the evaluation of AI-generated chest X‑ray reports. By quantifying the sensitivity of common metrics to these variations, the authors show that changes in reference reports can shift model rankings. They introduce a taxonomy of reporting variations and a rewriting method, ReRef, that preserves clinical meaning while altering style, and release a validated dataset of paired reference reports to aid future research.
arXiv:2609.01470v1 Announce Type: new Abstract: As AI systems are increasingly used to draft radiology reports, reliably evaluating their clinical quality remains a critical challenge. Large language...
The paper examines how Retrieval-Augmented Generation (RAG) and Named Entity Recognition (NER) affect the quality of lay summaries of radiology reports. Using a framework that extracts clinically relevant findings via NER and grounds them with RAG, the authors evaluate few‑shot and fine‑tuned versions of Qwen and BioBART. Results show that NER consistently improves readability and overall quality, RAG alone offers no benefit and can introduce hallucinations, and the best performance comes from fine‑tuned BioBART with NER.
arXiv:2609.22281v1 Announce Type: new Abstract: Foundation models have recently demonstrated strong capabilities across a wide range of medical imaging tasks. However, their performance in structured...