arXiv AI

No Free Lunch for Synthetic Images under Data Scarcity Conditions

arXiv:2606. 07640v1 Announce Type: cross Abstract: This study investigates the trade-offs between fidelity, privacy, and utility in synthetic data generation under conditions of data scarcity and privacy sensitivity.

arXiv Computer Vision
Sep 1

Coarse to Fine: Iterative Adversarial Neural Cellular Automata for Medical Image Synthesis

The paper introduces StyleGANCA, a lightweight neural cellular automata (NCA) based generative adversarial network designed for medical image synthesis. By combining a StyleGAN-inspired mapping network with adaptive style modulation in a multi-scale NCA framework, the model achieves high-quality image generation with far fewer parameters than existing adversarial, variational, diffusion, and NCA baselines. Experiments on BloodMNIST and PathMNIST show competitive FID and KID scores, and the synthetic images preserve class-specific information, effectively supporting downstream multi-class classifier training.

By Anh Thi Luu, Nick Lemke, Anirban Mukhopadhyay
Hugging Face Trending Papers
Jun 8

Synthetic but Not Realistic: The Evaluation Challenge in Generative Modelling for Structured Electronic Medical Records

Synthetic healthcare data are widely proposed as privacy-preserving substitutes for real patient data, yet their evaluation remains dominated by statistical similarity and predictive performance that do not reflect clinical validity. We introduce a multi-dimensional evaluation framework grounded in epidemiology, assessing descriptive fidelity, clinical utility, and structural validity, corresponding to descriptive, predictive, and causal questions.

arXiv Machine Learning
Jun 9

Synthetic but Not Realistic: The Evaluation Challenge in Generative Modelling for Structured Electronic Medical Records

arXiv:2606. 08903v1 Announce Type: new Abstract: Synthetic healthcare data are widely proposed as privacy-preserving substitutes for real patient data, yet their evaluation remains dominated by statistical similarity and predictive performance that do not reflect clinical validity.

By Nicholas I-Hsien Kuo, Blanca Gallego, Louisa Jorm
arXiv Machine Learning
Jul 27

The pretraining domain outweighs the training objective in setting the privacy-utility trade-off of differentially private medical image analysis

arXiv:2601. 19618v2 Announce Type: replace-cross Abstract: Differential privacy protects the patients whose images train medical imaging models, but it lowers diagnostic accuracy, and the initialization is the strongest known remedy.

By Soroosh Tayebi Arasteh, Mina Farajiamiri, Mahshad Lotfinia, Behrus Hinrichs-Puladi, Jonas Bienzeisler, Mohamed Alhaskir, Mirabela Rusu, Christiane Kuhl, Sven Nebelung, Daniel Truhn
Hugging Face Trending Papers
Jul 8

Collaborative Synthetic Data Generation for Knowledge Transfer in Federated Learning

One-shot federated learning (OSFL) addresses the communication overhead of federated learning by limiting training to a single round, but doing so without sacrificing model quality is non-trivial, particularly when client data distributions diverge. Recent work has addressed this challenge by aggregating client knowledge on the server through the construction of transferable synthetic datasets or distillates.

arXiv Machine Learning
Jun 26

Escaping Iterative Parameter-Space Noise: Differentially Private Learning with a Hypernetwork

arXiv:2606. 26772v1 Announce Type: new Abstract: Differentially private (DP) training of neural networks is often hindered by the large amount of noise required by gradient-based methods such as DP-SGD, which repeatedly inject high-dimensional noise in parameter space throughout training.

By Naoki Nishikawa, Shokichi Takakura, Satoshi Hasegawa
arXiv Computer Vision
Sep 4

Auditing Patient Privacy in Medical Generative Models: Scalable Memorization Detection with DeepSSIM++

DeepSSIM++ is a self‑supervised similarity metric designed to audit memorization in medical generative models at scale. It aggregates multi‑scale features and uses anatomy‑preserving augmentations to create an embedding space where cosine similarity approximates SSIM, removing the need for exact pixel‑level registration. Compared to existing baselines, DeepSSIM++ improves Macro F1 by 33–46 percentage points and speeds up large‑scale similarity computation by several orders of magnitude.

By Antonio Scardace, Francesco Guarnera, Sebastiano Battiato, Daniele Rav\`i