arXiv:2608. 03218v1 Announce Type: cross Abstract: Dataset distillation compresses a large training set into a compact synthetic set while retaining its downstream utility.
By Mingzhuo Li, Guang Li, Linfeng Ye, Jiafeng Mao, Takahiro Ogawa, Konstantinos N. Plataniotis, Miki Haseyama
The paper investigates whether pretrained image models can generalize to unseen datasets by clustering their embeddings. Using encoders trained only on ImageNet‑1k, both supervised and self‑supervised, the authors evaluate clustering performance on out‑of‑domain images. They find that supervised encoders perform better within the training domain, while self‑supervised encoders excel far outside it, and that fine‑tuning self‑supervised models reverses this trend. Additionally, the study shows that the silhouette score in UMAP‑reduced space correlates strongly with clustering accuracy, offering a proxy metric when labels are unavailable.
By Scott C. Lowe, Joakim Bruslund Haurum, Sageev Oore, Thomas B. Moeslund, Graham W. Taylor
DEFUSE is a backdoor detection framework for self‑supervised encoders that uses a conditional diffusion generative model to estimate representation‑conditioned image likelihoods. By fine‑tuning a pretrained diffusion model, DEFUSE performs semantic reconstruction in a reference encoder’s representation space, enabling it to detect backdoors without needing uninfected data or precomputed pseudo‑labels. Experiments show that DEFUSE outperforms existing detectors on both visual SSL and vision‑language encoders, reducing reliance on prior knowledge of the victim model or attack strategy.
By Tuo Chen, Jie Gui, Minjing Dong, Lanting Fang, Ju Jia, Benlei Cui, Jian Liu
arXiv:2407.03463v2 Announce Type: replace-cross
Abstract: In the realm of self-supervised learning (SSL), conventional wisdom has gravitated towards the utility of massive, general domain datasets fo...
By Jes\'us M Rodr\'iguez-de-Vera, Imanol G Estepa, Ignacio Saras\'ua, Bhalaji Nagarajan, Petia Radeva
The paper investigates the effect of using semantic positive pairs—different instances of the same class—in self‑supervised visual representation learning. By creating matched ImageNet‑1K subsets of augmented pairs and manually curated semantic pairs, the authors compare contrastive and non‑contrastive SSL methods under identical training conditions. Across transfer learning and object detection tasks, semantic‑pair pretraining consistently outperforms augmented‑pair pretraining, with contrastive methods like SimCLR showing the largest gains, indicating that semantic pairs foster additional invariances beyond standard augmentations.
By Mohammad Alkhalefi, Georgios Leontidis, Mingjun Zhong
AdaDim introduces a training strategy for self‑supervised learning that adaptively balances dimensionality increase and mutual information reduction. By gradually regularizing the projection head while encouraging feature decorrelation and sample uniformity, AdaDim achieves up to 3% performance gains over standard SSL baselines without relying on costly techniques such as queues or predictor networks. The method demonstrates that optimal SSL models do not simply maximize dimensionality or minimize mutual information, but find a trade‑off between the two.
By Kiran Kokilepersaud, Mohit Prabhushankar, Ghassan AlRegib