arXiv Machine Learning

3D-DLP: Self-Supervised 3D Object-Centric Scene Representation Learning

arXiv:2606. 19451v1 Announce Type: new Abstract: We introduce 3D-DLP, a self-supervised object-centric representation learning model that decomposes scene-level RGB-D or voxel observations into a set of 3D latent particles.

arXiv Computer Vision
Sep 18

ParticleSplat: Self-supervised Object-centric Latent Particle Splatting

ParticleSplat is a self‑supervised, object‑centric learning framework that extends the Deep Latent Particles (DLP) model into 3D by representing scenes as latent particles mapped to 3D Gaussian splats. It jointly encodes multiple camera views into a shared 3D latent space, enabling unsupervised learning of object masks and controllable 3D scene editing such as moving objects by manipulating latent particles. Experiments on simulated and real‑world datasets demonstrate that this 3D representation improves performance on downstream robotic manipulation tasks.

By Lyuxing He, Daniel Guo, Elizabeth Terveen, Deepak Pathak, David Held, Tal Daniel
arXiv Computer Vision
Sep 7

Object Concepts Emerge from Motion

The paper introduces a biologically inspired framework that learns object‑centric visual representations from raw videos without human annotations or camera calibration. By using motion boundaries detected via optical flow and clustering to create pseudo‑instance masks, the method supervises a single‑image encoder with pixel‑level pairwise metric learning. Training on 195 million pseudo‑labeled frames and expanding to 421 million frames through Motion‑Verified Self‑Training, the approach yields Swin‑based encoders that outperform or match supervised and self‑supervised baselines on tasks such as monocular depth estimation, 3D object detection, 3D occupancy prediction, and end‑to‑end planning.

By Boshi Li, Xiaohui Wang, Xiaoyang Wu, Zhichao Li, Ya Yang, Naiyan Wang
arXiv Computer Vision
Sep 18

FunArt: Decoding Functional Structure and Articulation from Generative 3D Latents

FunArt is a framework that builds articulation‑aware functional 3D scene graphs from a single static RGB‑D observation. It reconstructs object instances, converts their geometry into the O‑Voxel representation of TRELLIS.2, and uses a frozen sparse‑compression VAE as a structural prior. A lightweight query‑based decoder jointly segments movable parts and functional interactive elements while estimating motion type, axis, origin, and range, achieving state‑of‑the‑art results on the Articulate3D dataset.

By Dennis Rotondi, Abdelrhman Werby, Kai O. Arras