arXiv AI

Hypersolid: Emergent Vision Representations via Short-Range Repulsion

The paper introduces Hypersolid, a self‑supervised learning objective that uses short‑range repulsion to prevent representation collapse. It combines view alignment with local collision avoidance, creating a latent geometry of compact, semantically aligned neighborhoods with low anisotropy. This geometry improves unsupervised clustering and fine‑grained separation, though it reduces transferability.

arXiv AI
Sep 1

Social-JEPA: Emergent Geometric Isomorphism

arXiv:2603.02263v3 Announce Type: replace-cross Abstract: World models compress rich sensory streams into compact latent codes that anticipate future observations. We let separate agents acquire such...

By Haoran Zhang, Youjin Wang, Yi Duan, Rong Fu, Dianyu Zhao, Sicheng Fan, Shuaishuai Cao, Wentao Guo, Xiao Zhou
arXiv Computer Vision
Aug 28

DINOcular: Self-Supervised Visuospatial Representations

DINOcular is a self‑supervised framework that learns joint visuospatial representations from RGB‑D observations. It fuses depth‑derived geometric priors with a visual backbone using inter‑patch and intra‑patch fusion, allowing the model to encode both appearance and spatial structure efficiently. The resulting representation improves 3D awareness on multiple geometry benchmarks while staying competitive on standard RGB‑D semantic segmentation tasks.

By Farkhat Almukhamedov, Sami Azirar, Hermann Blum
arXiv AI
Jul 7

HASSL: Hierarchy-Aware Self-Supervised Learning Framework for Single Cell Microscopy

arXiv:2607. 04353v1 Announce Type: cross Abstract: Hierarchical structure is common in image data, where fine-grained clusters often merge into larger, coarser semantic groups.

By Julius Riel, Vishwa Mohan Singh, Sai Anirudh Aryasomayajula, Anuun Chinbat, Hannes Leonhard, Moritz Ladenburger, Frederik Alexander, Vishisht Choudhary, Fabio Laredo, Giacomo Masserdotti, Thorben Prein, Carsten Marr, Amirhossein Kardoost
arXiv AI
Jun 3

Anatomy-Anchored Self-Supervision: Distilling Vision Foundation Models for Invariant Ultrasound Representation

arXiv:2605. 25402v2 Announce Type: replace-cross Abstract: Self-supervised pre-training paradigm has gained increasing prominence for learning transferable representations in medical imaging, yet existing methods for ultrasound (US) images operate at the image or frame level, overlooking the anatomical context for clinical-aligned representation learning.

By Chunzheng Zhu, Yijun Wang, Jianxin Lin, Feng Wang, Hongwei Wang, Lei Zhao, Shengli Li, Kenli Li
arXiv Machine Learning
Jul 3

Object-centric LeJEPA

arXiv:2607. 02404v1 Announce Type: cross Abstract: Image encoders trained with LeJEPA can deliver strong features for downstream tasks, but, like other image-level self-supervised methods, typically require large training datasets.

By Jakob Geusen, Ender Konukoglu
Hugging Face Trending Papers
Aug 27

DINOcular: Self-Supervised Visuospatial Representations

DINOcular is a self‑supervised framework that learns joint visuospatial representations from RGB‑D data. It fuses depth‑derived geometric priors with a visual backbone using inter‑patch and intra‑patch techniques, allowing the model to encode both appearance and spatial structure efficiently. The resulting representation improves 3D awareness on multiple geometry benchmarks while staying competitive on standard RGB‑D semantic segmentation tasks.