Big, Bright, or Invisible: A Frozen-Feature Benchmark of 3D CT Foundation Models
arXiv:2608. 05960v1 Announce Type: cross Abstract: Routine CT interpretation is inherently comprehensive, capturing incidental findings across the entire scan volume.
arXiv:2607. 20993v1 Announce Type: cross Abstract: Large vision-language models are becoming increasingly dominant in 3D medical image interpretation, but we rarely know which internal units encode clinical findings or where that information lives in the representation.
arXiv:2608. 05960v1 Announce Type: cross Abstract: Routine CT interpretation is inherently comprehensive, capturing incidental findings across the entire scan volume.
The paper introduces SPAR‑Bench, a set of eight probes designed to test whether medical vision models can reason about anatomy in abdominal CT scans. Experiments across five architectures and three foundation models—both frozen and fine‑tuned—show that while models can recall canonical organ locations, they fail to perform relational reasoning or spatial comparisons within a patient, even under zero‑shot transfer. The study also demonstrates that pooled probing underestimates a model’s relational capabilities and that open‑weight multimodal large language models perform poorly on these tasks.
arXiv:2608. 08713v1 Announce Type: cross Abstract: Vision-language models offer a promising path toward automating radiology report generation, but applying them to full 3D CT volumes poses substantial computational challenges.
arXiv:2606. 03180v1 Announce Type: cross Abstract: Vision-language models (VLMs) for radiology have emerged as a scalable paradigm by leveraging image-report pairs naturally produced in clinical workflows.
arXiv:2607. 22771v1 Announce Type: cross Abstract: Picking the frozen image encoder for a 3D~CT vision--language model (VLM), together with the token-compression scheme on top of it, is a search over many candidates.
arXiv:2510. 15042v3 Announce Type: replace-cross Abstract: In the 3D medical image domain, vision-language pre-training is used to create vision-language encoders (VLEs) that can support radiologists by retrieving patients with similar abnormalities, predicting likelihoods of abnormality, or, with downstream adaptation, generating radiological reports.
DALE-CT introduces depth‑aware 2D slice encoders that learn an anatomical world model of chest CT scans without 3D or positional supervision. By sampling self‑supervised views across a physical $z$‑axis slab, the encoder captures how anatomy changes between neighboring slices, enabling it to recover slice ordering and distinguish slices by anatomy alone. The model, trained on a large 287k‑scan corpus, achieves state‑of‑the‑art performance on CT‑RATE and is released with full code and evaluation tools.
arXiv:2606. 28798v1 Announce Type: new Abstract: Objective: ICD codes are central to reimbursement, research, and population health surveillance, yet automated coding systems often struggle to integrate diagnostic signals from both clinical narratives and structured electronic health record (EHR) variables.
arXiv:2608. 12086v1 Announce Type: cross Abstract: Vision-language models, such as contrastive language-image pre-training (CLIP)-based approaches, have reached state-of-the-art (SOTA) results in medical artificial intelligence.
arXiv:2607. 25164v1 Announce Type: cross Abstract: A CT examination captures multiple organs, but many biomedical questions concern abnormalities, prognosis, or longitudinal change in a specific organ.
Integrating 3D medical images with vision-language models (VLMs) holds substantial promise for computer-aided diagnosis. However, volumetric images generate prohibitively long visual-token sequences with considerable spatial and inter-slice redundancy.
arXiv:2605.30984v2 Announce Type: replace-cross Abstract: Modern 3D medical vision-language models (VLMs) can generate fluent radiology-style text while exhibit critically low pathology detection and...