arXiv AI

HomeWorld: A Unified Floorplan-to-Furnished Framework for Generating Controllable, Densely Interactive Whole-Home Scenes

arXiv:2606. 06390v1 Announce Type: cross Abstract: Indoor scene generation is crucial for robot simulation and modern interior design.

Hugging Face Trending Papers
Aug 19

Beyond Placement and Articulation: Usage-Driven Code Scenes for Embodied Interaction

The paper introduces RoomWright, a code‑driven framework that generates 3D indoor scenes for embodied AI by focusing on functional usage rather than just visual layout. It performs usage‑driven object reasoning, treating anchors as task centers to select task‑required objects and their affordances, and compiles interactions into trigger‑condition‑effect rules that update object states. The system also addresses ambiguous object orientation through annotation‑guided usage cues, producing scenes that are executable, editable, and ready for simulation‑based policy learning.

arXiv Computer Vision
Sep 23

HARMONY: Hierarchical Agentic Reasoning for MONocular Image-to-Scene Synthesis

HARMONY is a hierarchical chain-of-thought framework that reconstructs complete 3D indoor scenes from a single monocular image. It combines agentic reasoning with visual geometry foundation models, starting with camera calibration and semantic layout recovery, then placing objects hierarchically while refining geometry with point cloud estimations. The method achieves semantically consistent scenes that align perceptually with the input image, outperforming existing baselines on synthetic and real-world data.

By Shufan Sun, Chen Wang, Enxin Song, Jiatao Gu, Lingjie Liu
arXiv Computer Vision
Aug 28

4DSynth: Controllable Procedural World Synthesis for Dynamic Embodied Simulation

4DSynth is a controllable procedural system that transforms natural-language descriptions, blueprint masks, or single photographs into editable 4D environments featuring explicit geometry, animated actors, collision-free trajectories, and physics-ready simulation states. The system unifies animation, camera planning, rendering, and task generation within a single geometry-grounded representation, enabling scalable creation of dynamic embodied simulation scenes. Using 4DSynth, the authors built 4DSynth-Nav, an interactive navigation benchmark that demonstrates the reproducibility and tunability of procedural failures across vision‑language models.

By Zehao Qi, Haochen Luo, Jia-Wang Bian, Zeyu Ma, Shuyang Sun
arXiv AI
Jul 14

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

arXiv:2607. 11643v1 Announce Type: cross Abstract: Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints.

By Xinghang Li, Jun Guo, Qiwei Li, Long Qian, Hang Lai, Yueze Wang, Hongyu Yan, Jiahang Cao, Xi Chen, Jingen Qu, Jiaxi Song, Nan Sun, Hanye Zhao, Futeng Liu, Wanli Peng, Heyun Wang, Yunhong Wang, Caoyu Xia, Jack Zhao, Diyun Xiang, Hangjun Ye, Heng Qu, Huaping Liu, Jason Li
arXiv Computation and Language
Sep 1

PlanCraft: Sketch, Refine, and Furnish for Architect-Inspired Progressive 3D Residential Scene Generation

PlanCraft introduces a progressive approach to 3D residential scene generation that mirrors how architects design: starting with rough sketches and refining them over time. It leverages a large dataset of real floor plans to train a SketchPlan module that generates partial sketches at various completion levels, a PlanCraft‑Diff module that sharpens these sketches into precise vector floor plans, and a PlanCraft‑Agent that furnishes rooms within the established spatial contract. The method outperforms existing 2D and 3D baselines, achieving a 61.1% lower FID and a 15‑point lead in expert‑rated spatial rationality, even with only 25% sketch completion.

By Pengyu Zeng, Yuqin Dai, Jun Yin, Ziyang Han, Ng Cheuk Hei, Jing Zhong, Chaoyang Shi, ZhanXiang Jin, Maowei Jiang, Shuai Lu
Hugging Face Trending Papers
Aug 27

4DSynth: Controllable Procedural World Synthesis for Dynamic Embodied Simulation

4DSynth is a controllable procedural system that transforms natural-language descriptions, blueprint masks, or single photographs into editable 4D environments featuring explicit geometry, animated actors, collision-free trajectories, and physics-ready simulation states. The system unifies animation, camera planning, rendering, and task generation within a single geometry-grounded representation, enabling scalable creation of diverse, interactive scenes. Using 4DSynth, the authors built 4DSynth-Nav, an interactive navigation benchmark that demonstrates the reproducibility of failures and tunable difficulty across three tiers for vision‑language models.