arXiv:2608.29519v1 Announce Type: new
Abstract: We introduce Function-Room Generation, a new indoor 3D scene generation setting that creates rooms supporting explicit functional goals rather than mer...
By Hao Feng, Zhi Zuo, MingJian Liang, Jingyu Hu, Xiaowei Hu, Liupengfei Wu, Dian Zhang, Guoxin Fang, Zhengzhe Liu
PlanCraft introduces a progressive approach to 3D residential scene generation that mirrors how architects design: starting with rough sketches and refining them over time. It leverages a large dataset of real floor plans to train a SketchPlan module that generates partial sketches at various completion levels, a PlanCraft‑Diff module that sharpens these sketches into precise vector floor plans, and a PlanCraft‑Agent that furnishes rooms within the established spatial contract. The method outperforms existing 2D and 3D baselines, achieving a 61.1% lower FID and a 15‑point lead in expert‑rated spatial rationality, even with only 25% sketch completion.
By Pengyu Zeng, Yuqin Dai, Jun Yin, Ziyang Han, Ng Cheuk Hei, Jing Zhong, Chaoyang Shi, ZhanXiang Jin, Maowei Jiang, Shuai Lu
arXiv:2607. 02407v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in 3D indoor synthesis for Manhattan environments.
By Xianhui Meng, Zirui Song, Yuchen Zhang, Li Zhang, Yongxuan Lv, Xiuying Chen, Kun Wang, Yan Luo, Kai Chen, Hangjun Ye, Long Chen, Jun Liu, Xiaoshuai Hao
arXiv:2606. 10953v1 Announce Type: new Abstract: Furnished floor plans are fundamental to real estate visualization, interior design, and architectural workflows.
By Fedor Rodionov, Aleksandar Cvejic, Michael Birsak, John Femiani, Peter Wonka
While traditional graphics methods often synthesize 3D indoor scenes autoregressively or hierarchically, recent vision-language model (VLM)-based generators predominantly adopt a one-shot paradigm where the full layout is planned at once. This one-shot approach often requires global re-optimization or complete reconstruction during interactive editing (e.
arXiv:2606. 05268v1 Announce Type: cross Abstract: We present a pipeline for building and aggregating task-specific, LLM-generated weak (imperfect) verifiers into a strong verifier for spatial layout domains.
By Sharon Zhang, R. Kenny Jones, Jiajun Wu, Maneesh Agrawala
arXiv:2609.36064v1 Announce Type: cross
Abstract: Foundation models are powerful generators, but many engineering domains require structured representations that general-purpose systems handle poorly...
By Sahand Rezaei-Shoshtari, Patryk Wozniczka, Shu Ishida, Gregg Streuber, Farnoosh Javadi, Jeffrey Landes, Angela Ju, Muhammad Azam, Bryan Lim, Johan Luttun, Indrajeet Haldar, Jonathan Shaw, Beatriz Guerra, Ivan Sosnovik, James Stoddart, Robert Giaquinto, Adam Gaier
arXiv:2512. 16275v2 Announce Type: replace-cross Abstract: Automated floor plan generation lies at the intersection of combinatorial search, geometric constraint satisfaction, and functional design requirements -- a confluence that has historically resisted a unified computational treatment.
By Mohamed Abouagour, Eleftherios Garyfallidis
arXiv:2608.03323v2 Announce Type: replace
Abstract: Estimating room layouts from multi-view imagery is a core task for indoor scene understanding. Existing methods are typically limited either by poo...
By Gustav Hanning, Shaohui Liu, R\'emi Pautrat, Marc Pollefeys, Kalle {\AA}str\"om, Viktor Larsson
arXiv:2609.23386v1 Announce Type: new
Abstract: Text-guided 3D building generation holds tremendous application potential, yet existing generative models typically output inseparable single meshes or...
By Xiang Tang, Ruotong Li, Xiaopeng Fan
ScenePilot introduces a retrieval‑augmented Grow‑and‑Repair framework for text‑driven 3D indoor scene generation. It uses a Hierarchical Retrieval‑Augmented Planning module to fetch room, group, and anchor layout priors, then incrementally inserts object groups with a base generator, while a Reinforcement Multimodal Repair module performs lightweight local corrections after each insertion and a final global repair. The approach is trained on a new SceneReverse‑17k dataset of perturbed scenes, enabling the policy to predict structured move‑rotate‑scale actions from rendered views, scene state, retrieved priors, and edit history, thereby improving physical plausibility, functional coherence, and controllability without heavy full‑scene optimization.
By Jiawei Zhang, Hongsong Wang, Pan Zhou
The paper introduces FactoSR, a factorized reinforcement learning framework designed to improve spatial reasoning in Vision‑Language Models by addressing a dimensional mismatch between 2D visual inputs and the 3D+temporal nature of the physical world. FactoSR decomposes the reasoning task into three orthogonal geometric sub‑objectives—planar correspondence (XY), depth consistency (Z), and temporal reversibility (T)—and optimizes these constraints within a unified policy learning mechanism. Experiments on multi‑view and video benchmarks show that this decomposition yields significant performance gains, achieving a 5.9% improvement on VSI‑Bench and 4.5% on All‑Angles‑Bench.
By Yijun Yang, Shenghe Zheng, Wenbo Li, Jianhui Liu, Haoze Sun, Yanbing Zhang, Jiaxiu Jiang, Lin Song, Haoyang Huang, Nan Duan, Lei Zhu