Building Rome from a Single Image
Single-image scene generation aims to produce a complete 3D scene mesh from a single image, including surfaces the camera did not observe. While pretrained 3D object generators encode a strong shape p...
Single-image scene generation aims to produce a complete 3D scene mesh from a single image, including surfaces the camera did not observe. While pretrained 3D object generators encode a strong shape p...
T3lescope is a generative surface reconstruction method that produces high‑fidelity 3D meshes from posed multi‑view images without per‑scene optimization. It uses a single fixed‑resolution generator in a coarse‑to‑fine cascade, where each level refines geometry within progressively finer spatial cells. Trained on cells at multiple scales, the model shares weights across all levels, allowing it to adapt the number of levels, cell scales, and locations at inference time and achieve consistent geometry across indoor, outdoor, and city‑scale scenes.
WorldSculpt presents a method for generating compositional 3D representations of cluttered scenes with hundreds of objects by adapting a single-object 3D generative prior to multi-view observations. The approach, built on Pixal3D with a multi-view conditioning pathway, can generalize to highly occluded scenes without scene-level training. The authors also introduce the UE-MeshyScene benchmark and demonstrate that their method outperforms prior approaches across various evaluation settings, including converting existing 3DGS worlds into compositional mesh scenes.
SceneReGen is a new framework for reconstructing 3D scenes from a single image by generating and assembling complete object meshes within a shared observation‑aligned scene frame. It uses selective pose factorization to encode each object’s observed orientation directly into the generated mesh, while estimating translation and scale from instance‑level and global scene cues. Evaluated on the 3D‑FUTURE dataset, SceneReGen outperforms existing methods on scene‑level metrics and shows strong performance on object‑level metrics, demonstrating its effectiveness in autonomous‑driving and embodied‑AI scenarios.
DistScene is a framework for generating 3D scenes from a single image by jointly modeling the environment and individual objects. It introduces Scene-Frame Generation to produce separate environment and object components in a shared coordinate frame, and Object-Centric Refinement to fine‑tune each object with scene context. The method also employs Object-to-Scene Distillation to transfer pretrained object-generation priors to scene generation using automatically composed synthetic scenes, achieving improved spatial coherence on indoor and outdoor benchmarks.
SpatialCrafter introduces a two‑stage framework for single‑image world modeling that first generates a global 3D proxy using a Point‑anchored Sparse Structure Flow module, then refines appearance with a Generative Deferred Refiner built on a video diffusion model. The method incorporates Parallel Geometry Injection and Proxy‑Aware Corruption training to integrate the proxy without disrupting the pretrained generative manifold, and it is evaluated on a newly constructed dataset of 115K scenes. Experiments demonstrate that SpatialCrafter outperforms existing approaches, reducing long‑term drift and maintaining consistency under rapid camera motion and extreme viewpoints.
arXiv:2609.23796v2 Announce Type: replace Abstract: Single-image 3D object generation can now produce high-fidelity assets, yet accurately placing them into a coherent scene layout remains an open ch...
arXiv:2609.23796v1 Announce Type: new Abstract: Single-image 3D object generation can now produce high-fidelity assets, yet accurately placing them into a coherent scene layout remains an open challe...
arXiv:2606. 08402v2 Announce Type: replace-cross Abstract: Generating complete 3D scenes from a single image requires inferring globally consistent geometry, object relationships, and environmental context from inherently ambiguous visual evidence.
arXiv:2606. 08402v1 Announce Type: cross Abstract: Generating complete 3D scenes from a single image requires inferring globally consistent geometry, object relationships, and environmental context from inherently ambiguous visual evidence.
arXiv:2609.10531v1 Announce Type: new Abstract: Image-to-3D models can generate visually compelling 3D assets from a single RGB image, but their geometry is often only loosely constrained by the avai...
DecomVoxel introduces a guided in‑situ denoising optimization that fuses 3D‑native priors with neural scene reconstruction to improve decompositional scene reconstruction. The method employs an epsilon‑based distillation loss for stable latent refinement and adaptive spatial guidance using occupied and vacant anchors with temporal annealing to reduce hallucinations and spatial drift. Experiments on Replica and ScanNet++ demonstrate that DecomVoxel outperforms state‑of‑the‑art approaches while preserving spatial layout, structural fidelity, and style‑consistent texture, yielding high‑quality textured meshes with clean topology.