From Alignment to Fusion in 3D Vision-Language
arXiv:2609.28222v1 Announce Type: new Abstract: Unified 3D vision-language systems must combine complementary geometry, scale, and appearance cues while supporting tasks from instance segmentation to...
PGL-3D introduces a progressive geometric learning framework for 3D visual query localization, where intermediate cuboids guide feature aggregation and refinement. The method predicts a complete cuboid for each proposal, selects reference observations via Query‑Tube‑Memory, pools query‑conditioned features, and re‑predicts refined cuboids. A training‑only objective, ST‑D9O, supervises cuboid geometry at every stage, yielding significant performance gains over prior baselines.
arXiv:2609.28222v1 Announce Type: new Abstract: Unified 3D vision-language systems must combine complementary geometry, scale, and appearance cues while supporting tasks from instance segmentation to...
arXiv:2608. 05774v1 Announce Type: cross Abstract: Joint-embedding predictive architectures learn by predicting latent representations of missing observations, yet many masked JEPAs are evaluated primarily through the encoders they produce.
arXiv:2608. 15710v1 Announce Type: cross Abstract: We address a fundamental gap in 3D-LLMs: existing models focus on single-object/scene description, struggling with detailed, inter-object comparison.
arXiv:2608.30618v1 Announce Type: new Abstract: Transformer-based decoders for 3D instance segmentation typically commit to a fixed number of queries and positional modeling calibrated on the trainin...
arXiv:2609.38620v1 Announce Type: new Abstract: Neural implicit representations have had a significant impact on scene reconstruction by enabling robots to build continuous, differentiable, and high-...
arXiv:2609.15639v1 Announce Type: new Abstract: Part-level control is essential for modern 3D asset creation, where objects are frequently edited, reused, animated, or fabricated through their indivi...
GAPrompt++ is a multi-granular geometry-aware prompting method designed to adapt pre-trained 3D vision models to downstream tasks efficiently. It introduces a Point Shift Prompter for multi-scale geometric feature extraction, a Keypoint Prompter for local geometric saliency, and a Prompt Propagation mechanism to embed these cues throughout the model hierarchy. Experiments demonstrate that GAPrompt++ outperforms other prompting-based PEFT methods and even surpasses full fine-tuning while using less than 2% trainable parameters, and the authors provide two new challenging benchmarks for future research.
UniQueR is a unified query‑based feedforward framework that reconstructs 3D scenes from unposed images by treating reconstruction as a sparse 3D query inference problem. It learns a compact set of 3D anchor points that serve as explicit geometric queries, allowing the network to infer scene structure—including occluded geometry—in a single forward pass. By encoding spatial and appearance priors directly in global 3D space and using a decoupled cross‑attention design, UniQueR achieves strong geometric expressiveness while reducing memory and computational cost, outperforming state‑of‑the‑art feedforward methods on Mip‑NeRF 360 and VR‑NeRF with far fewer primitives.
Recent advancements in 3D Gaussian Splatting (3DGS) have enabled language-guided scene understanding. However, existing Referring 3D Gaussian Splatting (R3DGS) methods are fundamentally restricted to single-target queries.
arXiv:2608.20720v1 Announce Type: new Abstract: Open-world 3D affordance grounding requires localizing functional object parts in 3D given free-form language queries. Existing methods typically assum...
arXiv:2606. 19733v1 Announce Type: cross Abstract: Efficiently retrieving specific 3D instances from large-scale scenes via natural language prompts remains a formidable challenge in multimedia analysis.
arXiv:2609.38177v1 Announce Type: cross Abstract: Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs...