ParticleSplat is a self‑supervised, object‑centric learning framework that extends the Deep Latent Particles (DLP) model into 3D by representing scenes as latent particles mapped to 3D Gaussian splats. It jointly encodes multiple camera views into a shared 3D latent space, enabling unsupervised learning of object masks and controllable 3D scene editing such as moving objects by manipulating latent particles. Experiments on simulated and real‑world datasets demonstrate that this 3D representation improves performance on downstream robotic manipulation tasks.
By Lyuxing He, Daniel Guo, Elizabeth Terveen, Deepak Pathak, David Held, Tal Daniel
We present ParticleSplat, a self-supervised object-centric learning method that decomposes scenes into a set of latent ''particles'' representing semantic entities through feedforward 3D Gaussian Spla...
arXiv:2503. 24009v3 Announce Type: replace-cross Abstract: Realistic simulation is critical for applications ranging from robotics to animation.
By Mikel Zhobro, Andreas Ren\'e Geist, Georg Martius
World models endow perceptual systems with the ability to predict how scenes evolve under interaction. They are most beneficial when trained on diverse volumes of data, to instill a rich prior into do...
arXiv:2609.19142v1 Announce Type: new
Abstract: World models endow perceptual systems with the ability to predict how scenes evolve under interaction. They are most beneficial when trained on diverse...
By Bardienus P. Duisterhof, Kaifeng Zhang, Adam Hung, Bowen Wen, Stan Birchfield, Yunzhu Li, Deva Ramanan, Jeffrey Ichnowski
arXiv:2604.19609v2 Announce Type: replace
Abstract: Transformers have become a common foundation across deep learning, yet 3D scene understanding still relies on specialized backbones with strong dom...
By Kadir Yilmaz, Adrian Kruse, Tristan H\"ofer, Daan de Geus, Bastian Leibe
World models enable agents to perform forward rollout and planning without real-world interaction. However, their application in open-world embodied intelligence remains limited by the high cost of action annotations and the heterogeneity of action spaces across platforms.
The paper introduces a biologically inspired framework that learns object‑centric visual representations from raw videos without human annotations or camera calibration. By using motion boundaries detected via optical flow and clustering to create pseudo‑instance masks, the method supervises a single‑image encoder with pixel‑level pairwise metric learning. Training on 195 million pseudo‑labeled frames and expanding to 421 million frames through Motion‑Verified Self‑Training, the approach yields Swin‑based encoders that outperform or match supervised and self‑supervised baselines on tasks such as monocular depth estimation, 3D object detection, 3D occupancy prediction, and end‑to‑end planning.
By Boshi Li, Xiaohui Wang, Xiaoyang Wu, Zhichao Li, Ya Yang, Naiyan Wang
FunArt is a framework that builds articulation‑aware functional 3D scene graphs from a single static RGB‑D observation. It reconstructs object instances, converts their geometry into the O‑Voxel representation of TRELLIS.2, and uses a frozen sparse‑compression VAE as a structural prior. A lightweight query‑based decoder jointly segments movable parts and functional interactive elements while estimating motion type, axis, origin, and range, achieving state‑of‑the‑art results on the Articulate3D dataset.
By Dennis Rotondi, Abdelrhman Werby, Kai O. Arras
arXiv:2503. 19947v2 Announce Type: replace-cross Abstract: Generalized metric depth understanding is critical for precise vision-guided robotics, which current state-of-the-art (SOTA) vision-encoders do not support.
By Paul Koch, J\"org Kr\"uger
arXiv:2609.38620v1 Announce Type: new
Abstract: Neural implicit representations have had a significant impact on scene reconstruction by enabling robots to build continuous, differentiable, and high-...
By Hanwen Cao, Wenqiang Wu, Kuang-Ting Tu, Mathias Otnes, Jeffrey Delmerico, Rui Wang, Yulun Tian, Nikolay Atanasov
arXiv:2411. 17790v3 Announce Type: replace-cross Abstract: Accurate 3D mapping in endoscopy enables quantitative, holistic lesion characterization within the gastrointestinal (GI) tract, requiring reliable depth and pose estimation.
By Ziang Xu, Bin Li, Yang Hu, Chenyu Zhang, James East, Sharib Ali, Jens Rittscher