arXiv Computer Vision

ParticleSplat: Self-supervised Object-centric Latent Particle Splatting

ParticleSplat is a self‑supervised, object‑centric learning framework that extends the Deep Latent Particles (DLP) model into 3D by representing scenes as latent particles mapped to 3D Gaussian splats. It jointly encodes multiple camera views into a shared 3D latent space, enabling unsupervised learning of object masks and controllable 3D scene editing such as moving objects by manipulating latent particles. Experiments on simulated and real‑world datasets demonstrate that this 3D representation improves performance on downstream robotic manipulation tasks.

arXiv Computer Vision
Sep 7

Object Concepts Emerge from Motion

The paper introduces a biologically inspired framework that learns object‑centric visual representations from raw videos without human annotations or camera calibration. By using motion boundaries detected via optical flow and clustering to create pseudo‑instance masks, the method supervises a single‑image encoder with pixel‑level pairwise metric learning. Training on 195 million pseudo‑labeled frames and expanding to 421 million frames through Motion‑Verified Self‑Training, the approach yields Swin‑based encoders that outperform or match supervised and self‑supervised baselines on tasks such as monocular depth estimation, 3D object detection, 3D occupancy prediction, and end‑to‑end planning.

By Boshi Li, Xiaohui Wang, Xiaoyang Wu, Zhichao Li, Ya Yang, Naiyan Wang
arXiv AI
Sep 25

KeyGen: Unsupervised Keypoint based Object-Centric Representations for Category-Level Policy Generalization

KeyGen is a framework that learns canonical 3D keypoints from point clouds to create structured, object‑centric representations for policy learning in robotic manipulation. By conditioning a visuomotor diffusion policy on these keypoints and object geometry, it predicts full manipulation trajectories that maintain geometric correspondence across different object instances. Experiments on a photorealistic simulation benchmark with three tasks show that KeyGen outperforms prior methods on both seen and unseen objects, scales with more demonstrations, remains robust to rescaling, and performs well in real‑world manipulation.

By Shuxin Cao, Liquan Wang, Masoud Moghani, Benjamin Joffe, Animesh Garg
arXiv Computer Vision
Sep 24

GaussianDS: Depth-supervised Semantic Gaussian Splatting for Scene Understanding

GaussianDS introduces a depth‑supervised framework for 3D Gaussian Splatting that jointly optimizes RGB appearance, depth, and compact semantics from scratch. By arranging multi‑view images into a pose‑aware pseudo‑video and propagating view‑consistent masks via SAM2, the method aligns semantic lifting with geometric cues, using depth supervision and edge‑aware refinement to curb semantic drift and boundary leakage. The approach achieves state‑of‑the‑art performance on LERF and 3D‑OVS benchmarks while preserving high‑fidelity reconstruction and enabling downstream tasks such as 3D object removal.

By Yufei Zhang, Chenlu Zhan, Hongwei Wang