arXiv:2608.29434v1 Announce Type: cross
Abstract: JEPA world models make latent-space planning a practical route to control, but they are built almost exclusively on images. Whether latent prediction...
By Fabio F. Oberweger, Michael Schwingshackl
arXiv:2607. 00514v1 Announce Type: cross Abstract: Automatic understanding of dynamic 4D point clouds, the 3D-point sequences captured over time by depth sensors and LiDAR, is central to robotics and embodied perception.
By Trung Thanh Nguyen, Hai Nguyen-Truong, Tu Vo, Hoang M. Truong, Tuan-Anh Vu
PGL-3D introduces a progressive geometric learning framework for 3D visual query localization, where intermediate cuboids guide feature aggregation and refinement. The method predicts a complete cuboid for each proposal, selects reference observations via Query‑Tube‑Memory, pools query‑conditioned features, and re‑predicts refined cuboids. A training‑only objective, ST‑D9O, supervises cuboid geometry at every stage, yielding significant performance gains over prior baselines.
By Liang Peng, Shizhuo Mu, Bohan Tan, Wenyuan Wang, Chen Zhao, Xingping Dong, Heng Fan, Libo Zhang, Bo Du
arXiv:2605.26949v2 Announce Type: replace
Abstract: 3D shape completion from partial scans remains challenging for unseen categories and noisy real-world observations, where geometry alone is often i...
By Furkan Mert Algan, Eckehard Steinbach
arXiv:2609.31595v1 Announce Type: new
Abstract: 3D scene graphs provide a structured representation of complex environments by encoding objects, their semantic attributes, and the spatial and functio...
By Luka Milivojevic, Nikola Popovic, Sayan Deb Sarkar, Sebastian Koch, Iro Armeni, Luc Van Gool, Danda Pani Paudel
arXiv:2608.24093v1 Announce Type: cross
Abstract: Self-supervised representation learning for 4D point cloud videos is challenging because annotations are costly and reconstruction-based pretraining...
By Jheng-Ling Lee, Shang-Tse Chen