arXiv AI

Learning an Interior Layout Policy in a Domain Specific Language Action Space

arXiv:2608. 07547v1 Announce Type: cross Abstract: Indoor scene layout generation is a challenging task in interior design.

arXiv Computation and Language
Sep 1

PlanCraft: Sketch, Refine, and Furnish for Architect-Inspired Progressive 3D Residential Scene Generation

PlanCraft introduces a progressive approach to 3D residential scene generation that mirrors how architects design: starting with rough sketches and refining them over time. It leverages a large dataset of real floor plans to train a SketchPlan module that generates partial sketches at various completion levels, a PlanCraft‑Diff module that sharpens these sketches into precise vector floor plans, and a PlanCraft‑Agent that furnishes rooms within the established spatial contract. The method outperforms existing 2D and 3D baselines, achieving a 61.1% lower FID and a 15‑point lead in expert‑rated spatial rationality, even with only 25% sketch completion.

By Pengyu Zeng, Yuqin Dai, Jun Yin, Ziyang Han, Ng Cheuk Hei, Jing Zhong, Chaoyang Shi, ZhanXiang Jin, Maowei Jiang, Shuai Lu
Hugging Face Trending Papers
Jul 15

ThinkBLOX: 3D Indoor Scene Generation with Progressive Reasoning

While traditional graphics methods often synthesize 3D indoor scenes autoregressively or hierarchically, recent vision-language model (VLM)-based generators predominantly adopt a one-shot paradigm where the full layout is planned at once. This one-shot approach often requires global re-optimization or complete reconstruction during interactive editing (e.

arXiv AI
4d ago

FLOORA: A Human-Aligned Domain-Specific Language Model for Architectural Design

arXiv:2609.36064v1 Announce Type: cross Abstract: Foundation models are powerful generators, but many engineering domains require structured representations that general-purpose systems handle poorly...

By Sahand Rezaei-Shoshtari, Patryk Wozniczka, Shu Ishida, Gregg Streuber, Farnoosh Javadi, Jeffrey Landes, Angela Ju, Muhammad Azam, Bryan Lim, Johan Luttun, Indrajeet Haldar, Jonathan Shaw, Beatriz Guerra, Ivan Sosnovik, James Stoddart, Robert Giaquinto, Adam Gaier
arXiv AI
Jul 28

GFLAN: Generative Functional Layouts

arXiv:2512. 16275v2 Announce Type: replace-cross Abstract: Automated floor plan generation lies at the intersection of combinatorial search, geometric constraint satisfaction, and functional design requirements -- a confluence that has historically resisted a unified computational treatment.

By Mohamed Abouagour, Eleftherios Garyfallidis
arXiv Computer Vision
Sep 17

PolyLayout: Multi-room Manhattan Layout Estimation

arXiv:2608.03323v2 Announce Type: replace Abstract: Estimating room layouts from multi-view imagery is a core task for indoor scene understanding. Existing methods are typically limited either by poo...

By Gustav Hanning, Shaohui Liu, R\'emi Pautrat, Marc Pollefeys, Kalle {\AA}str\"om, Viktor Larsson
arXiv AI
Sep 1

ScenePilot: Grow-and-Repair Policy for Text-Driven 3D Indoor Scene Generation

ScenePilot introduces a retrieval‑augmented Grow‑and‑Repair framework for text‑driven 3D indoor scene generation. It uses a Hierarchical Retrieval‑Augmented Planning module to fetch room, group, and anchor layout priors, then incrementally inserts object groups with a base generator, while a Reinforcement Multimodal Repair module performs lightweight local corrections after each insertion and a final global repair. The approach is trained on a new SceneReverse‑17k dataset of perturbed scenes, enabling the policy to predict structured move‑rotate‑scale actions from rendered views, scene state, retrieved priors, and edit history, thereby improving physical plausibility, functional coherence, and controllability without heavy full‑scene optimization.

By Jiawei Zhang, Hongsong Wang, Pan Zhou
arXiv Computer Vision
Sep 4

Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning

The paper introduces FactoSR, a factorized reinforcement learning framework designed to improve spatial reasoning in Vision‑Language Models by addressing a dimensional mismatch between 2D visual inputs and the 3D+temporal nature of the physical world. FactoSR decomposes the reasoning task into three orthogonal geometric sub‑objectives—planar correspondence (XY), depth consistency (Z), and temporal reversibility (T)—and optimizes these constraints within a unified policy learning mechanism. Experiments on multi‑view and video benchmarks show that this decomposition yields significant performance gains, achieving a 5.9% improvement on VSI‑Bench and 4.5% on All‑Angles‑Bench.

By Yijun Yang, Shenghe Zheng, Wenbo Li, Jianhui Liu, Haoze Sun, Yanbing Zhang, Jiaxiu Jiang, Lin Song, Haoyang Huang, Nan Duan, Lei Zhu