CheXGround: Anatomical Region Tokens for Grounded Longitudinal Chest X-ray Interpretation
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2608. 03890v1 Announce Type: cross Abstract: A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision thresholds, localize them spatially, and derive the anatomical measurements upon which many diagnoses depend.
In clinical practice, patients often undergo multiple imaging examinations over successive visits, yielding longitudinal data. Modeling such temporal information is crucial for reliable assessment of disease progression and treatment response.
Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level metrics that measure lexical overlap or aggregate clinical correctness. However, such metrics do not test whether individual diagnostic statements stem from the actual pathological evidence visible in the image.
arXiv:2607. 05880v1 Announce Type: cross Abstract: Imaging demand is growing faster than the radiology workforce can expand, and reporting backlogs cannot be resolved through training and recruitment alone.
The article reviews how multimodal large language models (MLLMs) are expanding radiology AI beyond image‑specific tasks to multimodal reasoning, yet volumetric radiology poses a representational challenge because clinical interpretation needs full 3‑D spatial context and quantitative data. It surveys over 200 studies, categorizing advances in volumetric representation, multimodal understanding, and agentic orchestration, and introduces a Claim‑Design‑Validation framework to align technical, workflow, and clinical claims. The review emphasizes that native volumetric modeling and agentic capabilities must match spatial, quantitative, contextual, and workflow demands, and that clinical credibility hinges on faithful 3‑D representation, traceable behavior, proper validation, and defined human oversight.
arXiv:2608. 00147v1 Announce Type: cross Abstract: Vision-language pretraining learns rich medical image representations from radiology reports, but previous model variants commonly operate within a single shared embedding space, so concept-level structure and interpretability must be recovered post hoc, limiting model transparency and, hence, clinical utility.