arXiv AI

Emergent Multi-View Geometry Through Self-Distillation

The paper introduces Poincar3, a self‑supervised method that learns multi‑view representations through self‑distillation rather than RGB reconstruction. By combining masked patch and image‑level distillation with a teacher that sees additional views, it trains from scratch without explicit 3D supervision. Poincar3 surpasses prior single‑ and multi‑view self‑supervised methods on tasks such as correspondence estimation, camera pose estimation, and 3D reconstruction, and its features encode camera motion more accurately thanks to a lightweight Poincaré adapter.

arXiv Computer Vision
Sep 7

Object Concepts Emerge from Motion

The paper introduces a biologically inspired framework that learns object‑centric visual representations from raw videos without human annotations or camera calibration. By using motion boundaries detected via optical flow and clustering to create pseudo‑instance masks, the method supervises a single‑image encoder with pixel‑level pairwise metric learning. Training on 195 million pseudo‑labeled frames and expanding to 421 million frames through Motion‑Verified Self‑Training, the approach yields Swin‑based encoders that outperform or match supervised and self‑supervised baselines on tasks such as monocular depth estimation, 3D object detection, 3D occupancy prediction, and end‑to‑end planning.

By Boshi Li, Xiaohui Wang, Xiaoyang Wu, Zhichao Li, Ya Yang, Naiyan Wang
arXiv Computer Vision
Aug 28

DINOcular: Self-Supervised Visuospatial Representations

DINOcular is a self‑supervised framework that learns joint visuospatial representations from RGB‑D observations. It fuses depth‑derived geometric priors with a visual backbone using inter‑patch and intra‑patch fusion, allowing the model to encode both appearance and spatial structure efficiently. The resulting representation improves 3D awareness on multiple geometry benchmarks while staying competitive on standard RGB‑D semantic segmentation tasks.

By Farkhat Almukhamedov, Sami Azirar, Hermann Blum
Hugging Face Trending Papers
Aug 27

DINOcular: Self-Supervised Visuospatial Representations

DINOcular is a self‑supervised framework that learns joint visuospatial representations from RGB‑D data. It fuses depth‑derived geometric priors with a visual backbone using inter‑patch and intra‑patch techniques, allowing the model to encode both appearance and spatial structure efficiently. The resulting representation improves 3D awareness on multiple geometry benchmarks while staying competitive on standard RGB‑D semantic segmentation tasks.

Hugging Face Trending Papers
Jul 21

IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer

Real-world spatial intelligence requires agents to understand scenes from continuous video streams, where objects move, persist, disappear, and reappear over time. While recent spatial foundation models have enabled generalizable feed-forward 3D reconstruction, most streaming methods remain geometry-centric and lack temporally consistent object-level understanding.

arXiv Computer Vision
Sep 18

ParticleSplat: Self-supervised Object-centric Latent Particle Splatting

ParticleSplat is a self‑supervised, object‑centric learning framework that extends the Deep Latent Particles (DLP) model into 3D by representing scenes as latent particles mapped to 3D Gaussian splats. It jointly encodes multiple camera views into a shared 3D latent space, enabling unsupervised learning of object masks and controllable 3D scene editing such as moving objects by manipulating latent particles. Experiments on simulated and real‑world datasets demonstrate that this 3D representation improves performance on downstream robotic manipulation tasks.

By Lyuxing He, Daniel Guo, Elizabeth Terveen, Deepak Pathak, David Held, Tal Daniel