Unified and Efficient Point-Line Local Features
arXiv:2608. 19894v1 Announce Type: new Abstract: Multi-view computer vision pipelines typically rely on accurate sparse keypoints and robust descriptors.
arXiv:2601. 18252v2 Announce Type: replace-cross Abstract: Wireframe parsing aims to recover line segments and their junctions to form a structured geometric representation useful for downstream tasks such as Simultaneous Localization and Mapping (SLAM).
arXiv:2608. 19894v1 Announce Type: new Abstract: Multi-view computer vision pipelines typically rely on accurate sparse keypoints and robust descriptors.
GALoc is a geometry-first approach for indoor localization that replaces depth prediction with gravity-aligned wireframes, ensuring verticality and coplanarity by construction. Using monocular RGB, camera intrinsics, relative poses, and IMU orientation, it builds a linear constraint matrix and finds the camera gauge that minimizes its smallest singular value through a global search. The resulting wireframes are projected into bird’s-eye-view layouts and matched against floorplans via a metric-free SE(2) search, achieving up to 88% sequential localization success on Gibson datasets and outperforming depth-based baselines when sufficient wall geometry is visible.
arXiv:2605.31577v2 Announce Type: replace Abstract: Recent feedforward 3D reconstruction methods predict point maps and estimate global 3D geometry remarkably well. However, their predictions still e...
GAPrompt++ is a multi-granular geometry-aware prompting method designed to adapt pre-trained 3D vision models to downstream tasks efficiently. It introduces a Point Shift Prompter for multi-scale geometric feature extraction, a Keypoint Prompter for local geometric saliency, and a Prompt Propagation mechanism to embed these cues throughout the model hierarchy. Experiments demonstrate that GAPrompt++ outperforms other prompting-based PEFT methods and even surpasses full fine-tuning while using less than 2% trainable parameters, and the authors provide two new challenging benchmarks for future research.
arXiv:2610.03717v1 Announce Type: cross Abstract: This paper examines the role of Novel View Synthesis (NVS) in geometric representation learning. In principle, NVS should reason about 3D scene struc...
The paper introduces Epipolar Distillation (EpiDistill), a method that transfers scale‑aware geometric priors from multi‑view models to monocular depth foundation models using Rectified Stereo Tokens. By preserving epipolar attention patterns, the single‑view model maintains geometric consistency without needing multi‑view inputs during inference. Experiments show significant improvements in zero‑shot metric depth estimation on challenging datasets such as ETH3D and DIODE, and the approach consistently boosts performance of state‑of‑the‑art ViT‑based models like UniDepthV2 and DepthPro.
arXiv:2604. 02546v2 Announce Type: replace-cross Abstract: Pretraining 3D encoders by aligning with Contrastive Language Image Pretraining (CLIP) has emerged as a promising direction to learn generalizable representations for 3D scene understanding.
arXiv:2512. 16919v2 Announce Type: replace-cross Abstract: Perceiving and reconstructing 3D scene geometry from visual inputs is crucial for autonomous driving.
arXiv:2608. 15284v1 Announce Type: cross Abstract: Navigation instruction generation from ego-centric RGB video in continuous environments is an important yet challenging task for human-robot interaction and scalable dataset construction.
arXiv:2606. 31148v1 Announce Type: cross Abstract: 3D Visual Grounding (3DVG) aims to localize target objects in 3D scenes given natural language descriptions.
arXiv:2603. 21046v2 Announce Type: replace-cross Abstract: UAVs play an important role in applications such as autonomous exploration, disaster response, and infrastructure inspection.
arXiv:2607.11498v2 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models require 3D spatial reasoning, yet RGB observations encode robot-object geometry only implicitly. Lifting...