arXiv:2608.22619v1 Announce Type: cross
Abstract: Generative segmentation provides an alternative to direct pixel-wise prediction by operating on learned latent representations, but effective image-t...
By Md Maklachur Rahman, Md Hasan Al Banna, Saraf Anjum, Mahmudul Hasan, Tracy Hammond
arXiv:2606. 17989v1 Announce Type: cross Abstract: Multi-contrast magnetic resonance imaging (MRI) provides complementary information for clinical diagnosis.
By Yonghao Chen, Sicheng Yang, Rui Tang, Lei Zhu
arXiv:2608. 08135v1 Announce Type: cross Abstract: Cross-modality medical image translation can reduce the burden of multi-modal acquisitions, yet the field remains constrained by two coupled limitations: methods operate on 2D slices or 3D patches rather than whole volumes, and train a separate model for each translation task.
By Daniele Molino, Alessio Zoboli, Camillo Maria Caruso, Valerio Guarrasi, Paolo Soda
The paper introduces a 3D-CLIP encoder trained with structured hard negatives to improve vision‑language alignment for text‑to‑CT generation. This encoder drives a latent diffusion model that operates directly in 3D latent space, eliminating spatial artifacts from super‑resolution pipelines. Experiments on the CT‑RATE dataset show state‑of‑the‑art image fidelity and factual correctness across 18 pathological conditions, with lower inference time and GPU memory usage than competing methods.
By Daniele Molino, Camillo Maria Caruso, Filippo Ruffini, Paolo Soda, Valerio Guarrasi
arXiv:2506. 00633v3 Announce Type: replace-cross Abstract: Generating semantically controllable 3D CT volumes from radiology reports requires more than a rich text encoder, it requires vision-language alignment grounded in volumetric space.
By Daniele Molino, Camillo Maria Caruso, Filippo Ruffini, Paolo Soda, Valerio Guarrasi
arXiv:2606. 19651v1 Announce Type: new Abstract: Three-dimensional (3D) brain MRI is central to clinical neurology and neuro-oncology, where generative models could augment under-represented cohorts, simulate disease trajectories, and support privacy-preserving data sharing.
By Max Van Puyvelde, Ibrahim Gulluk, Wim Van Criekinge, Olivier Gevaert
MedDiME is a latent-space, classifier‑guided diffusion framework designed for medical counterfactual image generation. It introduces a gradient‑driven adaptive masking mechanism that works directly in latent space, enabling spatially precise edits while avoiding the high computational and memory costs of pixel‑space methods. Experiments show MedDiME can produce high‑quality counterfactuals up to 40× faster and using 13× less GPU memory than previous diffusion baselines.
By Yan Zeng, Changlu Guo, Anders Nymark Christensen, Morten Rieger Hannemose, Anders Bjorholm Dahl
arXiv:2607. 02998v2 Announce Type: replace-cross Abstract: Controllable generative models of 3D medical images can synthesize volumes with specified clinical attributes, but this demands samples that are simultaneously high-fidelity, natively 3D, and faithful to the requested conditioning.
By Max Van Puyvelde, Halil Ibrahim Gulluk, Wim Van Criekinge, Olivier Gevaert
The paper introduces MR‑DiffuSR, a 3‑D latent diffusion framework that uses high‑resolution T1w structural priors to guide super‑resolution of thick‑slice FLAIR MRI scans. By applying cross‑modality structural swin attention and a mixed‑scale degradation strategy, the method avoids hallucinations and remains robust across varying slice thicknesses. On ADNI datasets, MR‑DiffuSR outperforms CNN and 2‑D diffusion baselines, achieving high PSNR, SSIM, and low LPIPS, and maintains strong white‑matter hyperintensity segmentation performance even at 7 mm equivalent slice thickness.
By Haoyu Lan, Jiazhen Zhang, John Onofrey, Bino Varghese, Nasim Sheikh-Bahaei, Arthur W. Toga, Jeiran Choupan
arXiv:2609.12860v1 Announce Type: new
Abstract: Computed tomography (CT) and positron emission tomography (PET) provide complementary anatomical and functional information for cancer diagnosis and tr...
By Sarita Mourya, Francesco Di Feola, Pierangelo Veltri, Paolo Soda
arXiv:2607. 02998v1 Announce Type: cross Abstract: Controllable generative models of 3D medical images can synthesize volumes with specified clinical attributes, but this demands samples that are simultaneously high-fidelity, natively 3D, and faithful to the requested conditioning.
By Max Van Puyvelde, Halil Ibrahim Gulluk, Wim Van Criekinge, Olivier Gevaert
arXiv:2608.28787v1 Announce Type: new
Abstract: Joint-embedding predictive architectures (JEPAs) have primarily been developed for self-supervised representation learning. Denoising JEPA (D-JEPA) rec...
By Meng Zhou, Wenhao You, Yuxing Chen, Yueying Tian