The paper introduces MedSegLatDiff, a diffusion-based framework that combines a variational autoencoder (VAE) with a latent diffusion model for medical image segmentation. By compressing images into a low-dimensional latent space, the method reduces noise and speeds up training, while a weighted cross‑entropy loss preserves tiny structures such as small nodules. Evaluated on ISIC‑2018, CVC‑Clinic, and LIDC‑IDRI datasets, MedSegLatDiff achieves state‑of‑the‑art Dice and IoU scores, generates diverse segmentation hypotheses, and produces confidence maps that enhance interpretability and reliability for clinical deployment.
By Ngoc Huynh Trinh, Hai Toan Nguyen, Son Ba Luong, Quoc Long Tran
We introduce TCAM-Diff, a novel 3D medical image generation model that reduces the memory requirements to encode and generate high-resolution 3D data. This model utilizes a decoder-only autoencoder method to learn triplane representation from dense volume and leverages generalization operations to prevent overfitting.
arXiv:2504. 19621v2 Announce Type: replace Abstract: Machine learning (ML) systems for medical imaging have demonstrated remarkable diagnostic capabilities, but their susceptibility to biases poses significant risks, since biases may negatively impact generalization performance.
By Haroui Ma, Francesco Quinzan, Theresa Willem, Stefan Bauer
arXiv:2606. 14759v1 Announce Type: cross Abstract: Cine cardiac magnetic resonance is the gold standard for assessing cardiac function, but the scarcity of public datasets limits the development of advanced data-driven models.
By Yiheng Cao (SyCoIA - IMT Mines Al\`es), Gustavo Andrade-Miranda (SyCoIA - IMT Mines Al\`es), Jiatian Zhang, Guillaume Sall\'e, Xin Gao
arXiv:2607. 02998v2 Announce Type: replace-cross Abstract: Controllable generative models of 3D medical images can synthesize volumes with specified clinical attributes, but this demands samples that are simultaneously high-fidelity, natively 3D, and faithful to the requested conditioning.
By Max Van Puyvelde, Halil Ibrahim Gulluk, Wim Van Criekinge, Olivier Gevaert
arXiv:2607. 05319v1 Announce Type: cross Abstract: We study why diffusion autoencoders can achieve similar image quality while learning substantially different latent structures.
By Rajat Rasal, Avinash Kori, Tian Xia, Ben Glocker
The paper introduces a 3D-CLIP encoder trained with structured hard negatives to improve vision‑language alignment for text‑to‑CT generation. This encoder drives a latent diffusion model that operates directly in 3D latent space, eliminating spatial artifacts from super‑resolution pipelines. Experiments on the CT‑RATE dataset show state‑of‑the‑art image fidelity and factual correctness across 18 pathological conditions, with lower inference time and GPU memory usage than competing methods.
By Daniele Molino, Camillo Maria Caruso, Filippo Ruffini, Paolo Soda, Valerio Guarrasi
arXiv:2605. 31162v1 Announce Type: cross Abstract: Unconditional diffusion models offer powerful generative priors, yet steering them toward aesthetically enhanced outputs remains largely unexplored.
By Shreyansh Modi, Akshat Tomar, Aarush Aggarwal
EraseSAE introduces a surgical concept erasure method for text-to-video diffusion models, using sparse autoencoders to decompose activations into interpretable, monosemantic features. The framework employs a contrastive attribution mechanism to isolate concept-specific kernels and applies timestep-resolved masks during inference to remove target concepts while preserving unrelated content. Experiments show that EraseSAE achieves precise, robust concept removal with minimal quality loss, outperforming existing methods.
arXiv:2607. 14580v1 Announce Type: cross Abstract: We present a novel system that integrates negative prompt optimization via a fine-tuned sequence-to-sequence LLM and latent-space classifier guidance to improve the quality of images generated by Stable Diffusion.
By Vaddi Charan Sai Nandan Reddy, Harini B, Chandana M S
arXiv:2608.22619v1 Announce Type: cross
Abstract: Generative segmentation provides an alternative to direct pixel-wise prediction by operating on learned latent representations, but effective image-t...
By Md Maklachur Rahman, Md Hasan Al Banna, Saraf Anjum, Mahmudul Hasan, Tracy Hammond
The paper presents a method for generating cardiac magnetic resonance (CMR) images conditioned on patient metadata using a pretrained latent diffusion model. By encoding structured clinical data and slice position as textual prompts and applying Metadata‑Free Classifier‑Free Guidance, Contrastive Batching, and Inverse‑Frequency Sampling, the authors improve the fidelity of synthetic images, achieving a 57% reduction in Fréchet Inception Distance compared to a baseline without these strategies. Evaluation on 59,058 UK Biobank CMR scans shows better distributional realism and subgroup alignment, though disease‑specific conditioning remains challenging.
By Marc Rodr\'iguez, Grzegorz Skorupko, Nay Aung, Steffen E Petersen, Karim Lekadir, Polyxeni Gkontra