Self-Supervised Representation-Guided Generative Dataset Distillation
arXiv:2608. 03218v1 Announce Type: cross Abstract: Dataset distillation compresses a large training set into a compact synthetic set while retaining its downstream utility.
arXiv:2606. 00514v1 Announce Type: new Abstract: Generative modeling and self-supervised representation learning (SSL) optimize structurally different objectives: generative training rewards distributional fidelity, while SSL rewards semantic coherence.
arXiv:2608. 03218v1 Announce Type: cross Abstract: Dataset distillation compresses a large training set into a compact synthetic set while retaining its downstream utility.
The paper investigates whether pretrained image models can generalize to unseen datasets by clustering their embeddings. Using encoders trained only on ImageNet‑1k, both supervised and self‑supervised, the authors evaluate clustering performance on out‑of‑domain images. They find that supervised encoders perform better within the training domain, while self‑supervised encoders excel far outside it, and that fine‑tuning self‑supervised models reverses this trend. Additionally, the study shows that the silhouette score in UMAP‑reduced space correlates strongly with clustering accuracy, offering a proxy metric when labels are unavailable.
DEFUSE is a backdoor detection framework for self‑supervised encoders that uses a conditional diffusion generative model to estimate representation‑conditioned image likelihoods. By fine‑tuning a pretrained diffusion model, DEFUSE performs semantic reconstruction in a reference encoder’s representation space, enabling it to detect backdoors without needing uninfected data or precomputed pseudo‑labels. Experiments show that DEFUSE outperforms existing detectors on both visual SSL and vision‑language encoders, reducing reliance on prior knowledge of the victim model or attack strategy.
arXiv:2407.03463v2 Announce Type: replace-cross Abstract: In the realm of self-supervised learning (SSL), conventional wisdom has gravitated towards the utility of massive, general domain datasets fo...
The paper investigates the effect of using semantic positive pairs—different instances of the same class—in self‑supervised visual representation learning. By creating matched ImageNet‑1K subsets of augmented pairs and manually curated semantic pairs, the authors compare contrastive and non‑contrastive SSL methods under identical training conditions. Across transfer learning and object detection tasks, semantic‑pair pretraining consistently outperforms augmented‑pair pretraining, with contrastive methods like SimCLR showing the largest gains, indicating that semantic pairs foster additional invariances beyond standard augmentations.
AdaDim introduces a training strategy for self‑supervised learning that adaptively balances dimensionality increase and mutual information reduction. By gradually regularizing the projection head while encouraging feature decorrelation and sample uniformity, AdaDim achieves up to 3% performance gains over standard SSL baselines without relying on costly techniques such as queues or predictor networks. The method demonstrates that optimal SSL models do not simply maximize dimensionality or minimize mutual information, but find a trade‑off between the two.
Dataset distillation compresses a large training set into a compact synthetic set while retaining its downstream utility. Most existing methods target randomly initialized networks, whereas modern vision systems often adapt frozen pretrained encoders with lightweight modules.
arXiv:2603. 15553v2 Announce Type: replace-cross Abstract: The landscape of self-supervised learning (SSL) is currently dominated by generative approaches (e.
Self-supervision is a powerful technique for learning visual representations from unlabeled data. Existing techniques primarily adopt a two-stage approach for self-supervised learning (SSL): a pretraining stage on unlabeled data followed by a finetuning stage on labeled data.
The paper reports a controlled study of self‑supervised learning (SSL) objectives for image and video pretraining under limited data, architecture, and compute budgets. It compares contrastive, reconstruction, feature‑prediction, and diffusion methods, finding that DINOv2‑style pretraining delivers the best overall performance. Combining DINOv2 with video SSL objectives such as VideoMAE improves image classification and segmentation but harms video tracking and camera‑pose estimation, highlighting a trade‑off between semantic and geometric learning.
arXiv:2609.36348v1 Announce Type: cross Abstract: Generative and representation learning remain asymmetrically connected: semantic representations are used to improve diffusion generation, whereas th...
The paper introduces Flow-Based Distribution Matching (FBDM), a non‑adversarial framework that learns self‑supervised representations using explicit geometric references and spherical conditional velocity regression. FBDM assigns augmented image views to shared target references while limiting reference usage, and employs an alignment loss to bring view representations closer. Experiments on datasets from CIFAR to ImageNet demonstrate that FBDM performs nearly as well as adversarial DM, outperforms existing SSL methods, and achieves a 1.48‑ to 1.83‑fold speedup with minimal GPU memory increase, while a theoretical analysis bounds downstream misclassification rates in terms of the pretraining loss.