arXiv AI

NeuSOGA3D: A Neuro-Symbolic Framework for Explainable 3D Geometric Reconstruction

NeuSOGA3D is a hybrid neuro‑symbolic framework that reconstructs 3D geometry from unorganized point clouds by combining learned perceptual priors with explicit symbolic geometric reasoning. It projects point clouds onto orthographic planes, builds symbolic implicit spline representations, and fuses them via shape‑preserving constructive solid geometry to produce a coarse visual hull. Additional detail is added through cross‑sectional decomposition and volumetric reconstruction with Partial Shape‑Preserving Splines, yielding CAD‑compatible, structurally meaningful models across all 40 ModelNet40 categories.

arXiv AI
Sep 2

Neuro-Symbolic Geometric Abstraction (NeuSOGA): From Observations to Symbolic Mathematical Representations

Neuro‑Symbolic Geometric Abstraction (NeuSOGA) is a framework that converts raw observations into explicit symbolic mathematical representations. It achieves this by sequentially generating topological and geometric abstractions, using tools such as Euclidean Distance Transforms, Segment Anything, and Implicit Area Splines. The resulting analytical implicit models are interpretable, editable, and support arbitrary‑order smoothness, additive composition, and closed‑form evaluation across diverse sensing modalities.

By Qingde Li, Qingqi Hong, Jie Tian
arXiv AI
Jun 2

Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs

arXiv:2606. 01215v1 Announce Type: cross Abstract: Current 3D spatial reasoning methods face a fundamental trade-off: neuro-symbolic 3D (NS3D) concept learners achieve interpretable reasoning through compositional programs but are constrained to closed-set concept vocabularies and simple programs; end-to-end 3D multi-modal LLMs (3D MLLMs) could handle complex natural language and open-vocabulary concepts but suffer from black-box reasoning without explicit spatial verification.

By Wentao Mo, Yang Liu
arXiv Computer Vision
2d ago

MEGA: Object-Level Mesh Extraction from 3D Gaussian Splatting via Spatial Visual Distillation

MEGA is a new framework that extracts object-level, watertight meshes from 3D Gaussian Splatting (3DGS) scenes. It uses a segment-then-mesh approach, leveraging Spatial Visual Distillation (SVD) to sample diverse camera views of each segmented object and train a mesh reconstruction model with photometric supervision. Experiments on popular benchmarks show that MEGA outperforms existing methods in accurately recovering object-level 3D occupancy and supports complex physical interactions by combining high-quality meshes with photorealistic 3DGS rendering.

By Liwei Liao, Yingkui Zhang, Qianqian Tong, Ronggang Wang
arXiv AI
Jul 24

3D-Aware VLMs with Implicit and Explicit Geometries

arXiv:2607. 21595v1 Announce Type: cross Abstract: Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial understanding and reasoning.

By Wenhao Li, Xueying Jiang, Quanhao Qian, Deli Zhao, Ran Xu, Shijian Lu, Gongjie Zhang
arXiv Computer Vision
Sep 17

Generalizable Neural Reconstruction of High-Fidelity Surfaces via Sparse Volumetric Representations

The paper introduces SVRecon, a generalizable neural surface reconstruction framework that uses sparse volumetric representations to achieve high-resolution 3D reconstruction. It employs a two-stage architecture: first predicting occupied voxels with an occupancy network, then rendering only within those regions using specialized sparse algorithms. This approach allows reconstruction at resolutions up to 512³ on 32 GB hardware, producing smoother and more precise surfaces, especially in sparse-view scenarios.

By Aoxiang Fan, Corentin Dumery, Nicolas Talabot, Ming Xu, Hieu Le, Pascal Fua
arXiv Computation and Language
4d ago

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

arXiv:2609.38177v1 Announce Type: cross Abstract: Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs...

By Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han, Mungyeom Kim, Minkyeong Jeon, Heeseong Shin, Wonjun Moon, Federico Tombari, Daniel Barath, Marc Pollefeys, Seungryong Kim, Sunghwan Hong
arXiv Computer Vision
Sep 18

GAPrompt++: Multi-Granular Geometry-Aware Point Cloud Prompt for 3D Vision Model

GAPrompt++ is a multi-granular geometry-aware prompting method designed to adapt pre-trained 3D vision models to downstream tasks efficiently. It introduces a Point Shift Prompter for multi-scale geometric feature extraction, a Keypoint Prompter for local geometric saliency, and a Prompt Propagation mechanism to embed these cues throughout the model hierarchy. Experiments demonstrate that GAPrompt++ outperforms other prompting-based PEFT methods and even surpasses full fine-tuning while using less than 2% trainable parameters, and the authors provide two new challenging benchmarks for future research.

By Zixiang Ai, Zhenyu Cui, Yufei Guo, Wenwen Qiang, Lei Chen, Jiwen Lu, Jiahuan Zhou
arXiv AI
Aug 11

P2Voxel: Pyramid Pivot Voxelization for 3D Mesh Tokenization

arXiv:2608. 07549v1 Announce Type: cross Abstract: Triangle meshes provide explicit and accurate surface geometry, yet their irregular topology connectivity makes 3D mesh tokenization a geometric sampling problem: how to sample and organize geometric evidence into compact, structured and learnable tokens.

By Zhenhong Sun, Haozhe Liu, Yifu Wang, Xibin Song, Senbo Wang, Huadong Mo, Daoyi Dong, Hongdong Li, Pan Ji
arXiv Computer Vision
Sep 23

Point Diffusion Mamba: Unified Diffusion-State-Space Modeling for Single-View 3D Reconstruction under Data Scarcity

Point Diffusion Mamba (PDM) is a new method that fuses diffusion models with state‑space modeling to perform single‑view 3D reconstruction when training data are scarce. It uses a lightweight reconstruction module for unordered point‑clouds, a Local Geometric Aggregation module combined with Mamba blocks to capture both global geometry and local detail, and a Hierarchical Feature Integration Network to merge high‑level semantic and local geometric features for each point. A Dynamic Weighted Sampling strategy further improves reconstruction quality by integrating generative priors, and experiments on ShapeNet and Pix3D show that PDM outperforms existing state‑of‑the‑art approaches.

By Wei Zhou, Xinzhe Shi, Xingxing Hao, Xing Hao, Kang Li, Jinye Peng, Ying He
arXiv Machine Learning
Jul 2

Efficient Compression of Structured and Unstructured Volumes via Learned 3D Gaussian Representation

arXiv:2607. 01164v1 Announce Type: new Abstract: Recent work has shown that implicit neural representations (INRs) can be trained to effectively compress structured and unstructured volume data, allowing for direct data querying with a reduced memory footprint.

By Landon Dyken, Sharmistha Chakrabarti, Nathan Debardeleben, Steve Petruzza, Qi Wu, Will Usher, Sidharth Kumar