PlanCraft introduces a progressive approach to 3D residential scene generation that mirrors how architects design: starting with rough sketches and refining them over time. It leverages a large dataset of real floor plans to train a SketchPlan module that generates partial sketches at various completion levels, a PlanCraft‑Diff module that sharpens these sketches into precise vector floor plans, and a PlanCraft‑Agent that furnishes rooms within the established spatial contract. The method outperforms existing 2D and 3D baselines, achieving a 61.1% lower FID and a 15‑point lead in expert‑rated spatial rationality, even with only 25% sketch completion.
By Pengyu Zeng, Yuqin Dai, Jun Yin, Ziyang Han, Ng Cheuk Hei, Jing Zhong, Chaoyang Shi, ZhanXiang Jin, Maowei Jiang, Shuai Lu
arXiv:2608.29519v1 Announce Type: new
Abstract: We introduce Function-Room Generation, a new indoor 3D scene generation setting that creates rooms supporting explicit functional goals rather than mer...
By Hao Feng, Zhi Zuo, MingJian Liang, Jingyu Hu, Xiaowei Hu, Liupengfei Wu, Dian Zhang, Guoxin Fang, Zhengzhe Liu
arXiv:2608. 07547v1 Announce Type: cross Abstract: Indoor scene layout generation is a challenging task in interior design.
By Yuhao Lu, Weichen Zhang, Wenyi Xiao, Haohui Chen, Yiyun Fei
The paper investigates how conditioned floor plan generation models perform when applied to datasets from different regions, revealing significant performance drops due to domain shift. To address this, the authors create a large synthetic training set that enforces physical constraints while deliberately reducing architectural realism, and show that pre‑training on this data boosts zero‑shot cross‑domain performance and speeds up fine‑tuning in low‑data scenarios.
By Matthieu Ospici, Arnaud Gueze, Luc Bourrat, Adrien Bernhardt
arXiv:2609.23386v1 Announce Type: new
Abstract: Text-guided 3D building generation holds tremendous application potential, yet existing generative models typically output inseparable single meshes or...
By Xiang Tang, Ruotong Li, Xiaopeng Fan
arXiv:2606. 10953v1 Announce Type: new Abstract: Furnished floor plans are fundamental to real estate visualization, interior design, and architectural workflows.
By Fedor Rodionov, Aleksandar Cvejic, Michael Birsak, John Femiani, Peter Wonka
arXiv:2609.36801v1 Announce Type: new
Abstract: Interactive simulations of embodied AI or spatial computing applications build on realistic 3D scenes that support daily activities. However, sparse, i...
By Minkwan Kim, Junho Kim, Seungmin Lee, Changwoon Choi, Young Min Kim
arXiv:2609.36064v1 Announce Type: cross
Abstract: Foundation models are powerful generators, but many engineering domains require structured representations that general-purpose systems handle poorly...
By Sahand Rezaei-Shoshtari, Patryk Wozniczka, Shu Ishida, Gregg Streuber, Farnoosh Javadi, Jeffrey Landes, Angela Ju, Muhammad Azam, Bryan Lim, Johan Luttun, Indrajeet Haldar, Jonathan Shaw, Beatriz Guerra, Ivan Sosnovik, James Stoddart, Robert Giaquinto, Adam Gaier
arXiv:2607. 16409v1 Announce Type: cross Abstract: Unified Multimodal Large Language Models (MLLMs) offer a promising paradigm for unifying visual understanding and generation, yet they still struggle to follow complex spatial instructions and logical constraints in controllable image generation.
By Junhao Liu, Jian-Wei Zhang, Tao Huang, Miles Yang, Zhao Zhong, Liefeng Bo
arXiv:2606. 06390v1 Announce Type: cross Abstract: Indoor scene generation is crucial for robot simulation and modern interior design.
By Wenbo Li, Xiaoliang Ju, Zipeng Qin, Rongyao Fang, Hongsheng Li
arXiv:2608.03323v2 Announce Type: replace
Abstract: Estimating room layouts from multi-view imagery is a core task for indoor scene understanding. Existing methods are typically limited either by poo...
By Gustav Hanning, Shaohui Liu, R\'emi Pautrat, Marc Pollefeys, Kalle {\AA}str\"om, Viktor Larsson
arXiv:2606. 08402v2 Announce Type: replace-cross Abstract: Generating complete 3D scenes from a single image requires inferring globally consistent geometry, object relationships, and environmental context from inherently ambiguous visual evidence.
By Jeonghwan Kim, Yushi Lan, Yongwei Chen, Hieu Trung Nguyen, Chuanyu Pan, Xingang Pan