arXiv Machine Learning By Zihan Zhou, Qifu Wen, Xi Zeng

SR-JEPA: Learning Predictive Latent State in 3D Scenes

Read the original on arXiv Machine Learning →

arXiv:2608. 05774v1 Announce Type: cross Abstract: Joint-embedding predictive architectures learn by predicting latent representations of missing observations, yet many masked JEPAs are evaluated primarily through the encoders they produce.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computer Vision
4d ago

PGL-3D: Towards Progressive Geometric Learning for 3D Visual Query Localization

PGL-3D introduces a progressive geometric learning framework for 3D visual query localization, where intermediate cuboids guide feature aggregation and refinement. The method predicts a complete cuboid for each proposal, selects reference observations via Query‑Tube‑Memory, pools query‑conditioned features, and re‑predicts refined cuboids. A training‑only objective, ST‑D9O, supervises cuboid geometry at every stage, yielding significant performance gains over prior baselines.

By Liang Peng, Shizhuo Mu, Bohan Tan, Wenyuan Wang, Chen Zhao, Xingping Dong, Heng Fan, Libo Zhang, Bo Du
arXiv Computer Vision
6d ago

GraphWrit3R: End-to-End 3D Scene Graph Writing

arXiv:2609.31595v1 Announce Type: new Abstract: 3D scene graphs provide a structured representation of complex environments by encoding objects, their semantic attributes, and the spatial and functio...

By Luka Milivojevic, Nikola Popovic, Sayan Deb Sarkar, Sebastian Koch, Iro Armeni, Luc Van Gool, Danda Pani Paudel