arXiv Computer Vision

PGL-3D: Towards Progressive Geometric Learning for 3D Visual Query Localization

PGL-3D introduces a progressive geometric learning framework for 3D visual query localization, where intermediate cuboids guide feature aggregation and refinement. The method predicts a complete cuboid for each proposal, selects reference observations via Query‑Tube‑Memory, pools query‑conditioned features, and re‑predicts refined cuboids. A training‑only objective, ST‑D9O, supervises cuboid geometry at every stage, yielding significant performance gains over prior baselines.

arXiv Machine Learning
Aug 7

SR-JEPA: Learning Predictive Latent State in 3D Scenes

arXiv:2608. 05774v1 Announce Type: cross Abstract: Joint-embedding predictive architectures learn by predicting latent representations of missing observations, yet many masked JEPAs are evaluated primarily through the encoders they produce.

By Zihan Zhou, Qifu Wen, Xi Zeng
arXiv Computer Vision
Sep 18

GAPrompt++: Multi-Granular Geometry-Aware Point Cloud Prompt for 3D Vision Model

GAPrompt++ is a multi-granular geometry-aware prompting method designed to adapt pre-trained 3D vision models to downstream tasks efficiently. It introduces a Point Shift Prompter for multi-scale geometric feature extraction, a Keypoint Prompter for local geometric saliency, and a Prompt Propagation mechanism to embed these cues throughout the model hierarchy. Experiments demonstrate that GAPrompt++ outperforms other prompting-based PEFT methods and even surpasses full fine-tuning while using less than 2% trainable parameters, and the authors provide two new challenging benchmarks for future research.

By Zixiang Ai, Zhenyu Cui, Yufei Guo, Wenwen Qiang, Lei Chen, Jiwen Lu, Jiahuan Zhou
arXiv AI
Sep 10

UniQueR: Unified Query-based Feedforward 3D Reconstruction

UniQueR is a unified query‑based feedforward framework that reconstructs 3D scenes from unposed images by treating reconstruction as a sparse 3D query inference problem. It learns a compact set of 3D anchor points that serve as explicit geometric queries, allowing the network to infer scene structure—including occluded geometry—in a single forward pass. By encoding spatial and appearance priors directly in global 3D space and using a decoupled cross‑attention design, UniQueR achieves strong geometric expressiveness while reducing memory and computational cost, outperforming state‑of‑the‑art feedforward methods on Mip‑NeRF 360 and VR‑NeRF with far fewer primitives.

By Chensheng Peng, Quentin Herau, Jiezhi Yang, Yichen Xie, Yihan Hu, Wenzhao Zheng, Matthew Strong, Masayoshi Tomizuka, Wei Zhan
arXiv Computation and Language
4d ago

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

arXiv:2609.38177v1 Announce Type: cross Abstract: Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs...

By Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han, Mungyeom Kim, Minkyeong Jeon, Heeseong Shin, Wonjun Moon, Federico Tombari, Daniel Barath, Marc Pollefeys, Seungryong Kim, Sunghwan Hong