arXiv Computer Vision By Liang Peng, Shizhuo Mu, Bohan Tan, Wenyuan Wang, Chen Zhao, Xingping Dong, Heng Fan, Libo Zhang, Bo Du

PGL-3D: Towards Progressive Geometric Learning for 3D Visual Query Localization

Read the original on arXiv Computer Vision →

PGL-3D introduces a progressive geometric learning framework for 3D visual query localization, where intermediate cuboids guide feature aggregation and refinement. The method predicts a complete cuboid for each proposal, selects reference observations via Query‑Tube‑Memory, pools query‑conditioned features, and re‑predicts refined cuboids. A training‑only objective, ST‑D9O, supervises cuboid geometry at every stage, yielding significant performance gains over prior baselines.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Machine Learning
Aug 7

SR-JEPA: Learning Predictive Latent State in 3D Scenes

arXiv:2608. 05774v1 Announce Type: cross Abstract: Joint-embedding predictive architectures learn by predicting latent representations of missing observations, yet many masked JEPAs are evaluated primarily through the encoders they produce.

By Zihan Zhou, Qifu Wen, Xi Zeng