arXiv Machine Learning By Liam Chalcroft

Tokenizer Generator Coupling in Medical Image Generation

Read the original on arXiv Machine Learning →

arXiv:2608. 07713v1 Announce Type: cross Abstract: Latent medical image generators usually treat the tokenizer as fixed preprocessing.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computer Vision
Sep 22

What Makes a Good Medical Image Tokenizer? Rethinking Reconstruction and Generation in Medical Image Tokenization

arXiv:2609.24691v1 Announce Type: new Abstract: Latent diffusion models now dominate medical image generation, and every such pipeline rests on a \emph{tokenizer} that compresses images into the late...

By Niklas Bubeck, Yundi Zhang, Vasiliki Sideri-Lampretsa, Julian McGinnis, Jiancheng Yang, Daniel Rueckert, Jiazhen Pan
arXiv AI
Aug 26

Metadata-Aware Adaptation of a Generative Foundation Model for Conditional CMR Synthesis

The paper presents a method for generating cardiac magnetic resonance (CMR) images conditioned on patient metadata using a pretrained latent diffusion model. By encoding structured clinical data and slice position as textual prompts and applying Metadata‑Free Classifier‑Free Guidance, Contrastive Batching, and Inverse‑Frequency Sampling, the authors improve the fidelity of synthetic images, achieving a 57% reduction in Fréchet Inception Distance compared to a baseline without these strategies. Evaluation on 59,058 UK Biobank CMR scans shows better distributional realism and subgroup alignment, though disease‑specific conditioning remains challenging.

By Marc Rodr\'iguez, Grzegorz Skorupko, Nay Aung, Steffen E Petersen, Karim Lekadir, Polyxeni Gkontra
arXiv AI
Jul 23

Self-supervision drives representational convergence in medical foundation models more than clinical supervision

arXiv:2607. 20274v1 Announce Type: cross Abstract: Medical image encoders from different groups are increasingly treated as interchangeable, on the assumption that scale and clinical supervision concentrate their representations onto a shared structure.

By Soroosh Tayebi Arasteh, Sebastian Ziegelmayer, Mahshad Lotfinia, Lisa Adams, Sven Nebelung, Jakob Nikolas Kather, Daniel Truhn
arXiv Machine Learning
1d ago

The Null Is the Hard Part: Exact Tests for Memorization in Generative Models

The paper critiques current memorization audits for generative models, arguing that lacking a proper null distribution leads to misleading conclusions. It introduces two exact null tests—one permutation test for whole models and a calibrated test for single images—showing that many previously flagged memorizations disappear under these stricter controls. The authors also propose a scale‑restricted statistic based on the Intersection Euler Characteristic Profile to better detect distinct copied images.

By Sushovan Majhi, Pramita Bagchi