arXiv Computer Vision

Scene Retargeting: Learning Object Placement with Analogical Transfer

arXiv AI
Sep 1

ScenePilot: Grow-and-Repair Policy for Text-Driven 3D Indoor Scene Generation

ScenePilot introduces a retrieval‑augmented Grow‑and‑Repair framework for text‑driven 3D indoor scene generation. It uses a Hierarchical Retrieval‑Augmented Planning module to fetch room, group, and anchor layout priors, then incrementally inserts object groups with a base generator, while a Reinforcement Multimodal Repair module performs lightweight local corrections after each insertion and a final global repair. The approach is trained on a new SceneReverse‑17k dataset of perturbed scenes, enabling the policy to predict structured move‑rotate‑scale actions from rendered views, scene state, retrieved priors, and edit history, thereby improving physical plausibility, functional coherence, and controllability without heavy full‑scene optimization.

By Jiawei Zhang, Hongsong Wang, Pan Zhou
Hugging Face Trending Papers
Jul 15

ThinkBLOX: 3D Indoor Scene Generation with Progressive Reasoning

While traditional graphics methods often synthesize 3D indoor scenes autoregressively or hierarchically, recent vision-language model (VLM)-based generators predominantly adopt a one-shot paradigm where the full layout is planned at once. This one-shot approach often requires global re-optimization or complete reconstruction during interactive editing (e.

Hugging Face Trending Papers
Aug 19

Beyond Placement and Articulation: Usage-Driven Code Scenes for Embodied Interaction

The paper introduces RoomWright, a code‑driven framework that generates 3D indoor scenes for embodied AI by focusing on functional usage rather than just visual layout. It performs usage‑driven object reasoning, treating anchors as task centers to select task‑required objects and their affordances, and compiles interactions into trigger‑condition‑effect rules that update object states. The system also addresses ambiguous object orientation through annotation‑guided usage cues, producing scenes that are executable, editable, and ready for simulation‑based policy learning.

arXiv Computer Vision
Sep 22

Mitigating Domain Shift in Conditioned Floor Plan Generation: Synthetic Pre-training for Data-Efficient Adaptation

The paper investigates how conditioned floor plan generation models perform when applied to datasets from different regions, revealing significant performance drops due to domain shift. To address this, the authors create a large synthetic training set that enforces physical constraints while deliberately reducing architectural realism, and show that pre‑training on this data boosts zero‑shot cross‑domain performance and speeds up fine‑tuning in low‑data scenarios.

By Matthieu Ospici, Arnaud Gueze, Luc Bourrat, Adrien Bernhardt
arXiv Computation and Language
Sep 1

PlanCraft: Sketch, Refine, and Furnish for Architect-Inspired Progressive 3D Residential Scene Generation

PlanCraft introduces a progressive approach to 3D residential scene generation that mirrors how architects design: starting with rough sketches and refining them over time. It leverages a large dataset of real floor plans to train a SketchPlan module that generates partial sketches at various completion levels, a PlanCraft‑Diff module that sharpens these sketches into precise vector floor plans, and a PlanCraft‑Agent that furnishes rooms within the established spatial contract. The method outperforms existing 2D and 3D baselines, achieving a 61.1% lower FID and a 15‑point lead in expert‑rated spatial rationality, even with only 25% sketch completion.

By Pengyu Zeng, Yuqin Dai, Jun Yin, Ziyang Han, Ng Cheuk Hei, Jing Zhong, Chaoyang Shi, ZhanXiang Jin, Maowei Jiang, Shuai Lu
arXiv Computer Vision
Sep 23

Fysiverse-3D-Vision Technical Report: Generating Executable 3D Worlds from Images through Unified Spatial Reasoning

Fysiverse-3D-Vision is a unified vision‑language‑geometry framework that reconstructs executable 3D scenes from a single image. It separates spatial layout reasoning from asset synthesis, using a shared representation where spatial reasoning and geometric reconstruction reinforce each other. The model employs a Transformer that integrates textual supervision, semantic visual cues, and geometric representations, and includes an object‑conditioned layout module to predict object translation, rotation, and scale while maintaining physical consistency through collision‑aware optimization.

By Dingkang Yang, Yizhou Liu, Wendong Cheng, Zizhi Chen, Shunli Wang, Yang Liu, Hongsheng Li, Lihua Zhang