arXiv AI
Sep 2

ViTAMINS: An Empirical Study of Training Self-Supervised Vision Transformers with Synthetic Hard Negatives

ViTAMINS is a method that incorporates synthetic hard negatives into unsupervised vision transformer pretraining to enhance representation quality. The approach is evaluated on ImageNet and a range of downstream tasks—including transfer learning, image retrieval, copy detection, and image/video segmentation—showing significant performance gains. The synthetic negatives also lead to emergent properties, such as representations that encode explicit semantic information and act as strong classifiers, improving over baselines by up to 11.3%.

By Nikos Giakoumoglou, Andreas Floros, Kleanthis-Marios Papadopoulos, Tania Stathaki
arXiv AI
Aug 24

StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models

StateSight is a new benchmark designed to isolate and evaluate the ability of vision‑language models to reconstruct latent spatial structure from a single image. It consists of three procedurally generated task families—cube‑net opposite‑face reasoning, occluded cube‑tower counting, and 4‑neighbor connected‑component counting—each with 300 deterministic prompts and exact‑match scoring. The benchmark also includes a companion dataset, StateSight‑Steps, with 900 image‑text examples and 3,600 intermediate visual states to aid analysis of reconstruction errors.

By Michelle Lin