arXiv Computer Vision

ReconPlusGen: Injecting Reconstruction Prior into Multi-view 3D Generation through Noise Inversion and Modulation

ReconPlusGen introduces a method that injects a reconstruction prior into multi‑view 3D generation. By predicting a point cloud in canonical space from multiple input images, the method deterministically injects the geometry into a diffusion process via noise inversion and then modulates the noise to maintain generative flexibility for completing unseen areas and refining visible geometry. The paper presents qualitative results on benchmark and real‑world images, along with an illustration of the reconstruction‑guided noise initialization and modulation.

Hugging Face Trending Papers
Sep 10

ReconPlusGen: Injecting Reconstruction Prior into Multi-view 3D Generation through Noise Inversion and Modulation

The paper introduces ReconPlusGen, a method that injects a reconstruction prior into multi‑view 3D generation. It predicts a point cloud in canonical space from multiple input images, then deterministically injects this geometry into a diffusion process via noise inversion and modulates the noise to maintain generative flexibility for completing unseen areas and refining visible geometry. Qualitative results show reconstruction on benchmark and real‑world images, along with an illustration of the reconstruction‑guided noise initialization and modulation.

arXiv Computer Vision
Aug 24

RecGen3D: Reconstruction-Guided 3D Generation in a Shared Canonical Space

RecGen3D is a framework that merges feed‑forward reconstruction and diffusion‑based generation to address the trade‑off between reconstruction fidelity and generative plausibility in sparse‑view 3D modeling. By aligning both models in a shared canonical space and using decoupled cooperative learning, the system stabilizes training and allows the reconstruction module to supply canonical geometric anchors while the diffusion generator refines and completes the structure. Experiments show that RecGen3D outperforms existing methods in producing complete and consistent 3D models from sparse observations.

By Zhisheng Huang, Jiahao Chen, Cheng Lin, Chenyu Hu, Hanzhuo Huang, Zhengming Yu, Mengfei Li, Yuheng Liu, Zekai Gu, Zibo Zhao, Yuan Liu, Xin Li, Wenping Wang
arXiv Computer Vision
Sep 23

Point Diffusion Mamba: Unified Diffusion-State-Space Modeling for Single-View 3D Reconstruction under Data Scarcity

Point Diffusion Mamba (PDM) is a new method that fuses diffusion models with state‑space modeling to perform single‑view 3D reconstruction when training data are scarce. It uses a lightweight reconstruction module for unordered point‑clouds, a Local Geometric Aggregation module combined with Mamba blocks to capture both global geometry and local detail, and a Hierarchical Feature Integration Network to merge high‑level semantic and local geometric features for each point. A Dynamic Weighted Sampling strategy further improves reconstruction quality by integrating generative priors, and experiments on ShapeNet and Pix3D show that PDM outperforms existing state‑of‑the‑art approaches.

By Wei Zhou, Xinzhe Shi, Xingxing Hao, Xing Hao, Kang Li, Jinye Peng, Ying He
arXiv Computer Vision
Sep 3

RoGe: Novel View Synthesis via End-to-End Implicit Reconstruction and Generation

RoGe is a new end‑to‑end framework for novel view synthesis that jointly learns an implicit 3D scene representation and a video diffusion model. It eliminates the need for explicit 3D intermediates by querying the implicit scene with camera rays to produce geometric features that condition the diffusion model. Experiments on DL3DV show that RoGe surpasses reconstruction‑based, generation‑based, and hybrid baselines in image quality and temporal consistency, and ablations confirm the benefits of ray‑queried features and joint training.

By Xiaolei Lang, Ze Kang, Zehao Huang, Naiyan Wang
arXiv Computer Vision
3d ago

T3lescope: Arbitrary-Resolution High-Fidelity Generative Surface Reconstruction from Images

T3lescope is a generative surface reconstruction method that produces high‑fidelity 3D meshes from posed multi‑view images without per‑scene optimization. It uses a single fixed‑resolution generator in a coarse‑to‑fine cascade, where each level refines geometry within progressively finer spatial cells. Trained on cells at multiple scales, the model shares weights across all levels, allowing it to adapt the number of levels, cell scales, and locations at inference time and achieve consistent geometry across indoor, outdoor, and city‑scale scenes.

By Atsuhiro Noguchi, Tianhan Xu, Yiming Liang, Yuta Kikuchi, Masahiro Ishiyama, Shintaro Takagi, Hitoshi Murai, Eiichi Matsumoto