arXiv Machine Learning

Semi-Supervised Conditional Diffusion via Label Augmentation

arXiv:2607. 16685v1 Announce Type: cross Abstract: Conditional diffusion models have become a powerful and flexible framework for learning complex conditional distributions from labeled data.

arXiv AI
Jun 10

MMD Guidance: Training-Free Distribution Adaptation for Diffusion Models via Maximum Mean Discrepancy Guidance

arXiv:2601. 08379v2 Announce Type: replace-cross Abstract: Pre-trained diffusion models have emerged as powerful generative priors for both unconditional and conditional sample generation, yet their outputs often deviate from the characteristics of user-specific target data.

By Matina Mahdizadeh Sani, Nima Jamali, Mohammad Jalali, Farzan Farnia
arXiv AI
Sep 3

No Data Wasted: A Semi-supervised Generative Model for Incomplete Multi-view Data Integration with Missing Labels

The paper presents a semi‑supervised generative model for multi‑view learning that handles missing views and missing labels. It combines a likelihood‑based approach for unlabeled data with an information bottleneck (IB) framework for labeled data, incorporating modality‑specific information and cross‑view mutual information maximization to learn a shared latent space. Experiments show improved predictive and generative performance on complex datasets with limited labeled samples.

By Yiyang Shen, Weiran Wang
arXiv Computer Vision
Sep 2

Semi-Supervised Biomedical Image Segmentation via Diffusion Models and Teacher-Student Co-Training

The paper presents a semi‑supervised biomedical image segmentation method that uses a diffusion‑based teacher–student framework. The teacher is pretrained via unsupervised diffusion reconstruction and then co‑trained with a student, leveraging supervised labels and cross pseudo‑supervision on unlabeled data. A multi‑round extension generates multiple stochastic reconstructions to further refine pseudo‑labels, achieving competitive or superior results on several 2D and 3D biomedical datasets, especially when labels are scarce.

By Luca Ciampi, Gabriele Lagani, Giuseppe Amato, Fabrizio Falchi
arXiv Machine Learning
Aug 26

Generalization, memorization, and overfitting for diffusion models trained in the lazy high-dimensional regime

The paper investigates diffusion models trained in a lazy high‑dimensional regime, extending benign overfitting theory to generative settings. By analyzing denoising score matching in a vector‑valued RKHS with an inner‑product kernel, the authors derive exact risk trajectories under gradient flow when the number of samples scales proportionally with dimensionality. These trajectories reveal three distinct phases—spectral generalization, noise‑dominated interpolation, and empirical Bayes memorization—whose interplay shapes the distribution of generated samples.

By Hugo Latourelle-Vigeant, Sinho Chewi, Aram-Alexandre Pooladian, John Sous, Theodor Misiakiewicz
arXiv Machine Learning
Aug 31

Diffusion models as plug-and-play priors

The paper explores using denoising diffusion generative models as plug‑and‑play priors for high‑dimensional inference problems. By combining a pre‑trained diffusion prior with a differentiable auxiliary constraint, the authors enable approximate inference through iterative differentiation across multiple noisy versions of the data. This framework opens possibilities for conditional generation, image segmentation, and novel combinatorial optimization algorithms.

By Alexandros Graikos, Esmeralda S. Whitammer, Nebojsa Jojic, Dimitris Samaras
arXiv Machine Learning
Jun 5

Inverse Entropic Optimal Transport Solves Semi-supervised Learning via Data Likelihood Maximization

arXiv:2410. 02628v5 Announce Type: replace Abstract: Learning conditional distributions $\pi^*(\cdot|x)$ is a central problem in machine learning, which is typically approached via supervised methods with paired data $(x,y) \sim \pi^*$.

By Mikhail Persiianov, Arip Asadulaev, Nikita Andreev, Nikita Starodubcev, Dmitry Baranchuk, Anastasis Kratsios, Evgeny Burnaev, Alexander Korotin