arXiv Machine Learning By Changyu Liu, Yuling Jiao, Jian Huang

Semi-Supervised Conditional Generative Learning through Stochastic Interpolation and Sufficient Representations

Read the original on arXiv Machine Learning →

arXiv:2607. 16725v1 Announce Type: cross Abstract: Conditional generative modeling remains a challenging problem in semi-supervised settings where labeled data is scarce but unlabeled samples are abundant.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 3

No Data Wasted: A Semi-supervised Generative Model for Incomplete Multi-view Data Integration with Missing Labels

The paper presents a semi‑supervised generative model for multi‑view learning that handles missing views and missing labels. It combines a likelihood‑based approach for unlabeled data with an information bottleneck (IB) framework for labeled data, incorporating modality‑specific information and cross‑view mutual information maximization to learn a shared latent space. Experiments show improved predictive and generative performance on complex datasets with limited labeled samples.

By Yiyang Shen, Weiran Wang
arXiv Machine Learning
Aug 27

JEPAMatch: Geometric Representation Shaping for Semi-Supervised Learning

JEPAMatch introduces a new semi‑supervised learning framework that replaces traditional output‑thresholding with explicit geometric shaping of latent representations. By combining the FlexMatch loss with a latent‑space regularization inspired by LeJEPA, the method encourages isotropic Gaussian structure in the embedding space, mitigating class imbalance and noisy pseudo‑labels. Experiments on CIFAR‑100, STL‑10, and Tiny‑ImageNet show consistent performance gains and faster convergence compared to existing FixMatch‑based baselines.

By Ali Aghababaei-Harandi, Aude Sportisse, Massih-Reza Amini
arXiv Machine Learning
Sep 25

Sufficiently Reduced Distributional Regression

Sufficiently Reduced Distributional Regression (SRDR) is a generative approach that merges conditional distribution estimation with nonlinear sufficient dimension reduction (SDR). By framing SDR as a risk minimization problem using strictly proper scoring rules, SRDR jointly learns a dimension reduction map and a generative prediction model through minimization of the energy score, which can be estimated via sampling. The method extends to multi‑environment data and classification, and theoretical results show convergence of estimated conditional distributions in energy distance, implying asymptotic sufficiency. In experiments on CT slice localization, superconductivity data, and digit classification, SRDR recovers low‑dimensional sufficient structure and matches or surpasses state‑of‑the‑art nonlinear SDR methods in representation quality and predictive performance.

By Alexander Henzi, Tiange Liu, Xinwei Shen