LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
HARMONY is a hierarchical chain-of-thought framework that reconstructs complete 3D indoor scenes from a single monocular image. It combines agentic reasoning with visual geometry foundation models, starting with camera calibration and semantic layout recovery, then placing objects hierarchically while refining geometry with point cloud estimations. The method achieves semantically consistent scenes that align perceptually with the input image, outperforming existing baselines on synthetic and real-world data.
arXiv:2605. 10873v2 Announce Type: replace-cross Abstract: Recovering editable CAD programs from images or 3D observations is central to AI-assisted design, but progress is difficult to measure because existing evaluations are fragmented across datasets, modalities, and metrics.
WorldSculpt presents a method for generating compositional 3D representations of cluttered scenes with hundreds of objects by adapting a single-object 3D generative prior to multi-view observations. The approach, built on Pixal3D with a multi-view conditioning pathway, can generalize to highly occluded scenes without scene-level training. The authors also introduce the UE-MeshyScene benchmark and demonstrate that their method outperforms prior approaches across various evaluation settings, including converting existing 3DGS worlds into compositional mesh scenes.
arXiv:2606. 08402v2 Announce Type: replace-cross Abstract: Generating complete 3D scenes from a single image requires inferring globally consistent geometry, object relationships, and environmental context from inherently ambiguous visual evidence.
Code4Scene is a benchmark that evaluates coding agents on constructing and editing 3D scenes in Unreal Engine. It tests agents on two tasks: construction, where they must build a scene from open‑ended language, and editing, where they must recover a target scene from reference images while preserving everything else. The benchmark measures task fulfillment, artifact integrity, and physical validity, revealing that construction and editing performance are correlated but not interchangeable, with agents struggling most with spatial composition and precise edits.
arXiv:2606. 08402v1 Announce Type: cross Abstract: Generating complete 3D scenes from a single image requires inferring globally consistent geometry, object relationships, and environmental context from inherently ambiguous visual evidence.