Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments
arXiv:2607. 02407v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in 3D indoor synthesis for Manhattan environments.
arXiv:2502. 06819v2 Announce Type: replace Abstract: This paper presents a framework for generating 3D indoor scenes from text prompts.
arXiv:2607. 02407v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in 3D indoor synthesis for Manhattan environments.
arXiv:2606. 08402v2 Announce Type: replace-cross Abstract: Generating complete 3D scenes from a single image requires inferring globally consistent geometry, object relationships, and environmental context from inherently ambiguous visual evidence.
arXiv:2606. 08402v1 Announce Type: cross Abstract: Generating complete 3D scenes from a single image requires inferring globally consistent geometry, object relationships, and environmental context from inherently ambiguous visual evidence.
The paper introduces LEGO, a benchmark dataset pairing user text descriptions with human‑annotated fine‑grained constraints and reference 3D scenes, and LEGO‑Eval, an evaluation framework that decomposes descriptions into atomic constraints and verifies each using grounding and spatial reasoning tools. It demonstrates that LEGO‑Eval detects misalignment more accurately than existing methods and that current 3D scene synthesis approaches achieve at most a 10% success rate on this benchmark.
arXiv:2601.14056v2 Announce Type: replace-cross Abstract: Training robust visual surveillance models requires large-scale datasets with precise spatial annotations, yet collecting real surveillance d...
arXiv:2606. 06390v1 Announce Type: cross Abstract: Indoor scene generation is crucial for robot simulation and modern interior design.
arXiv:2603. 16085v2 Announce Type: replace-cross Abstract: Recent breakthroughs in 3D generation have enabled the synthesis of high-fidelity individual assets.
ScenePilot introduces a retrieval‑augmented Grow‑and‑Repair framework for text‑driven 3D indoor scene generation. It uses a Hierarchical Retrieval‑Augmented Planning module to fetch room, group, and anchor layout priors, then incrementally inserts object groups with a base generator, while a Reinforcement Multimodal Repair module performs lightweight local corrections after each insertion and a final global repair. The approach is trained on a new SceneReverse‑17k dataset of perturbed scenes, enabling the policy to predict structured move‑rotate‑scale actions from rendered views, scene state, retrieved priors, and edit history, thereby improving physical plausibility, functional coherence, and controllability without heavy full‑scene optimization.
arXiv:2609.23386v1 Announce Type: new Abstract: Text-guided 3D building generation holds tremendous application potential, yet existing generative models typically output inseparable single meshes or...
arXiv:2606. 24206v1 Announce Type: cross Abstract: Recent breakthroughs in 3D generation have advanced notably with the development of text-to-image diffusion model.
WorldSculpt presents a method for generating compositional 3D representations of cluttered scenes with hundreds of objects by adapting a single-object 3D generative prior to multi-view observations. The approach, built on Pixal3D with a multi-view conditioning pathway, can generalize to highly occluded scenes without scene-level training. The authors also introduce the UE-MeshyScene benchmark and demonstrate that their method outperforms prior approaches across various evaluation settings, including converting existing 3DGS worlds into compositional mesh scenes.
The paper introduces RoomWright, a code‑driven framework that generates 3D indoor scenes for embodied AI by focusing on functional usage rather than just visual layout. It performs usage‑driven object reasoning, treating anchors as task centers to select task‑required objects and their affordances, and compiles interactions into trigger‑condition‑effect rules that update object states. The system also addresses ambiguous object orientation through annotation‑guided usage cues, producing scenes that are executable, editable, and ready for simulation‑based policy learning.