Learning Where to Look: Anatomical Grounding and Guided Attention for Cardiac MRI Vision-Language Models
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2509.22404v2 Announce Type: replace Abstract: Anatomical understanding, which is the ability to identify, localize, or segment anatomical structures, is critical in medical image analysis; howe...
arXiv:2608. 03890v1 Announce Type: cross Abstract: A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision thresholds, localize them spatially, and derive the anatomical measurements upon which many diagnoses depend.
arXiv:2608.30758v1 Announce Type: new Abstract: Recent radiology multi-modal language models have made substantial progress in chest X-ray report generation, visual question answering, and temporal r...
arXiv:2609.24057v1 Announce Type: cross Abstract: Medical image interpretation is central to diagnosis and care, yet adapting general-purpose multimodal large language models (MLLMs) often requires r...
arXiv:2608.30975v1 Announce Type: cross Abstract: Cardiac magnetic resonance imaging (CMR) produces rich sequential data such as temporal cine videos and spatial LGE/mapping stacks, yet most deep lea...
arXiv:2609.32352v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have shown increasing potential for medical image understanding, yet their capabilities in ophthalmic imaging r...