Tokenizer Generator Coupling in Medical Image Generation
arXiv:2608. 07713v1 Announce Type: cross Abstract: Latent medical image generators usually treat the tokenizer as fixed preprocessing.
arXiv:2606. 09601v1 Announce Type: new Abstract: Conditional generators provide a natural tool for controllable generation, including settings where the desired condition is a new composition of observed attributes or experimental factors.
arXiv:2608. 07713v1 Announce Type: cross Abstract: Latent medical image generators usually treat the tokenizer as fixed preprocessing.
arXiv:2607. 02637v1 Announce Type: cross Abstract: Recent generative models can produce high-quality synthetic images, offering scalable training training data for data-hungry models.
arXiv:2512.17730v2 Announce Type: replace Abstract: Detectors of AI-generated images tend to inherit the biases of the data they are trained on: models fitted to GAN imagery learn to treat GAN-specif...
arXiv:2607. 14984v1 Announce Type: new Abstract: Per-subgroup fairness audits of medical image classifiers face a sample-size problem: minority subgroups in held-out test sets have so few samples that the resulting confidence intervals on per-subgroup performance are wider than the bias the audit is meant to detect.
arXiv:2609.14124v1 Announce Type: cross Abstract: Medical image analysis is often hindered by biased datasets, which can lead to biased models and limited clinical applicability. A promising strategy...
arXiv:2607. 18088v1 Announce Type: new Abstract: Standard evaluation of many recognition systems contains distribution shift by construction, since benchmarks place disjoint conditions in the training and test splits.
arXiv:2605.10894v2 Announce Type: replace Abstract: Deep learning models in medical imaging often fail when deployed in new clinical environments due to distribution shifts in demographics, scanner h...
The paper introduces the Real‑Calibrated Synthetic‑First Data Engine, a modular pipeline that integrates controllable diffusion‑based synthetic image generation with multi‑stage curation, filtering, and optional uncertainty‑driven selection and human verification. Designed as a CLI‑based framework, it allows independent configuration of generation, filtering, selection, and validation modules to enhance reproducibility and flexibility in real‑world data workflows. Empirical tests on human pose estimation demonstrate that synthetic data can boost a real‑data baseline when used as low‑cost augmentation, though synthetic‑only training still lags behind real‑only performance, underscoring the importance of data‑centric orchestration in low‑data regimes.
The paper introduces a label‑free method called AURCC for selecting the best foundational model for medical image classification when the target domain lacks labels. AURCC uses a pseudo‑label discrepancy computed by the SUDO framework to score models without fine‑tuning. Experiments on chest X‑ray data across three inter‑hospital shifts show that AURCC closely matches the true model ranking, outperforming simple source‑accuracy baselines especially when source data are limited.
arXiv:2608. 03990v1 Announce Type: new Abstract: Synthetic histopathology image generation has emerged as an approach that may address data scarcity in computational pathology, yet current evaluation methodologies may not fully assess synthetic data quality for medical applications.
The paper introduces a classifier‑free method for generating visual counterfactual explanations (VCEs) using Contrastive Analysis (CA). By separating generative factors common to two datasets from those specific to each class, the approach swaps only the salient factors to produce counterfactual images, thereby avoiding reliance on classifier decision boundaries. Leveraging StyleGAN2’s high‑quality synthesis and a feature‑space latent representation, the method supports multiple salient factors per dataset and achieves superior counterfactual quality on three medical imaging datasets.
arXiv:2603. 16551v2 Announce Type: replace-cross Abstract: Generative models are increasingly used to augment medical imaging datasets for fairer AI, yet a key assumption often goes unexamined: that generators produce equally high-quality images across demographic groups.