Self-Supervised Representation-Guided Generative Dataset Distillation
arXiv:2608. 03218v1 Announce Type: cross Abstract: Dataset distillation compresses a large training set into a compact synthetic set while retaining its downstream utility.
arXiv:2608. 03218v1 Announce Type: cross Abstract: Dataset distillation compresses a large training set into a compact synthetic set while retaining its downstream utility.
Dataset distillation compresses a large training set into a compact synthetic set while retaining its downstream utility. Most existing methods target randomly initialized networks, whereas modern vision systems often adapt frozen pretrained encoders with lightweight modules.
arXiv:2602. 07345v2 Announce Type: replace-cross Abstract: Distribution Matching Distillation (DMD) is a powerful acceleration paradigm, yet its stability is often compromised in Forbidden Zone, regions where the real teacher provides unreliable guidance while the fake teacher exerts insufficient repulsive force.
The paper introduces Flow-Based Distribution Matching (FBDM), a non‑adversarial method that learns self‑supervised representations by aligning images to explicit geometric references through spherical conditional velocity regression. By using an ETF‑inspired reference, FBDM allows more reference components than the flow dimension while maintaining geometric separation, and it incorporates an alignment loss to bring augmented views closer together. Experiments on datasets from CIFAR to ImageNet show that FBDM performs nearly as well as adversarial distribution‑matching methods, achieves a 1.48‑ to 1.83‑fold speedup, and offers a theoretical bound on downstream misclassification rates.
arXiv:2407.03463v2 Announce Type: replace-cross Abstract: In the realm of self-supervised learning (SSL), conventional wisdom has gravitated towards the utility of massive, general domain datasets fo...
arXiv:2609.36348v1 Announce Type: cross Abstract: Generative and representation learning remain asymmetrically connected: semantic representations are used to improve diffusion generation, whereas th...
The paper introduces Flow-Based Distribution Matching (FBDM), a non‑adversarial framework that learns self‑supervised representations using explicit geometric references and spherical conditional velocity regression. FBDM assigns augmented image views to shared target references while limiting reference usage, and employs an alignment loss to bring view representations closer. Experiments on datasets from CIFAR to ImageNet demonstrate that FBDM performs nearly as well as adversarial DM, outperforms existing SSL methods, and achieves a 1.48‑ to 1.83‑fold speedup with minimal GPU memory increase, while a theoretical analysis bounds downstream misclassification rates in terms of the pretraining loss.
arXiv:2609.36099v1 Announce Type: new Abstract: Multi-step matching models, including flow and diffusion models, produce high-quality outputs but incur substantial inference costs and may reproduce u...
arXiv:2607. 14703v1 Announce Type: cross Abstract: Multiple instance learning (MIL) has become the main paradigm for whole-slide image (WSI) analysis in computational pathology.
arXiv:2603. 25144v2 Announce Type: replace-cross Abstract: Dataset distillation (DD) compresses a large training set into a small synthetic set, reducing storage and training cost, and has shown strong results on general benchmarks.
arXiv:2609.40333v1 Announce Type: new Abstract: Self-supervised learning draws inspiration from infant visual development, yet standard training pipelines bear little resemblance to it: images are in...
The paper introduces the Real‑Calibrated Synthetic‑First Data Engine, a modular pipeline that integrates controllable diffusion‑based synthetic image generation with multi‑stage curation, filtering, and optional uncertainty‑driven selection and human verification. Designed as a CLI‑based framework, it allows independent configuration of generation, filtering, selection, and validation modules to enhance reproducibility and flexibility in real‑world data workflows. Empirical tests on human pose estimation demonstrate that synthetic data can boost a real‑data baseline when used as low‑cost augmentation, though synthetic‑only training still lags behind real‑only performance, underscoring the importance of data‑centric orchestration in low‑data regimes.