LLM agents are increasingly used to translate natural language into 3D scenes in a procedural way, but existing systems focus on static output. Dynamic 4D scenes from text alone, in which liquids flow, particles emit, rigid bodies cascade, and articulated mechanisms move, remain largely unexplored despite their value as editable content and as physics-grounded training data for video generation and embodied AI.
arXiv:2607. 01766v1 Announce Type: new Abstract: LLM agents are increasingly used to translate natural language into 3D scenes in a procedural way, but existing systems focus on static output.
By Chunjiang Liu, Xiaoyuan Wang, Haoyu Chen, Yizhou Zhao, Ming-Hsuan Yang, L\'aszl\'o A. Jeni
The paper introduces LEGO, a benchmark dataset pairing user text descriptions with human‑annotated fine‑grained constraints and reference 3D scenes, and LEGO‑Eval, an evaluation framework that decomposes descriptions into atomic constraints and verifies each using grounding and spatial reasoning tools. It demonstrates that LEGO‑Eval detects misalignment more accurately than existing methods and that current 3D scene synthesis approaches achieve at most a 10% success rate on this benchmark.
By Minseok Kang, Dongwook Choi, Gyeom Hwangbo, Seungwon Lim, Kai Tzu-iunn Ong, Jinyoung Yeo
arXiv:2607. 27380v2 Announce Type: replace-cross Abstract: Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt.
By Haodong Li, Tianfei Ren, Xiaoxiao Ma, Chunmei Qing, Zhen Fang, Sipeng He, Ziyu Guo, Haoyu Wu, Juanxi Tian, Yihang Zou, Ruichuan An, Dongzhi Jiang, Boxue Yang, Ji Xie, Xu Huang, Wenhao Yan, Jialv Zou, Zhengrong Yue, Yaxin Luo, Xiaotong Li, Yuzhu Wang, Junyan Ye, Jinjing Zhao, Zehui Chen, Lin Chen, Renye Yan, Feng Zhao, Pheng-Ann Heng
arXiv:2602. 10840v2 Announce Type: replace Abstract: Large language models (LLMs) have been widely studied in areas such as mathematical reasoning, complex coding, and scientific problem solving.
By Yanan Wang, Renxi Wang, Yongxin Wang, Xuezhi Liang, Fajri Koto, Timothy Baldwin, Xiaodan Liang, Haonan Li
4DSynth is a controllable procedural system that transforms natural-language descriptions, blueprint masks, or single photographs into editable 4D environments featuring explicit geometry, animated actors, collision-free trajectories, and physics-ready simulation states. The system unifies animation, camera planning, rendering, and task generation within a single geometry-grounded representation, enabling scalable creation of dynamic embodied simulation scenes. Using 4DSynth, the authors built 4DSynth-Nav, an interactive navigation benchmark that demonstrates the reproducibility and tunability of procedural failures across vision‑language models.
By Zehao Qi, Haochen Luo, Jia-Wang Bian, Zeyu Ma, Shuyang Sun
arXiv:2603. 07053v3 Announce Type: replace Abstract: Scientists face significant visualization challenges as time-varying datasets grow in speed and volume, often requiring specialized infrastructure and expertise to handle massive datasets.
By Ishrat Jahan Eliza, Xuan Huang, Aashish Panta, Alper Sahistan, Zhimin Li, Amy A. Gooch, Valerio Pascucci
4DSynth is a controllable procedural system that transforms natural-language descriptions, blueprint masks, or single photographs into editable 4D environments featuring explicit geometry, animated actors, collision-free trajectories, and physics-ready simulation states. The system unifies animation, camera planning, rendering, and task generation within a single geometry-grounded representation, enabling scalable creation of diverse, interactive scenes. Using 4DSynth, the authors built 4DSynth-Nav, an interactive navigation benchmark that demonstrates the reproducibility of failures and tunable difficulty across three tiers for vision‑language models.
arXiv:2609.10540v1 Announce Type: new
Abstract: Recent video world models generate increasingly realistic and interactive visual experiences, yet lack reliable mechanisms for maintaining persistent w...
By Zheng-Hui Huang, Guixu Lin, Jiacheng Lin, Yi-Chuan Huang, Ruihan Yu, Muyao Niu, Siqi Yang, Yu-Lun Liu, Yung-Yu Chuang, Kaipeng Zhang, Zhixiang Wang
arXiv:2502. 06819v2 Announce Type: replace Abstract: This paper presents a framework for generating 3D indoor scenes from text prompts.
By Yao Wei, Matteo Toso, Pietro Morerio, Changjae Oh, Michael Ying Yang, Alessio Del Bue
The paper introduces Physically Plausible Video Generation (PPVG), a method that generates videos consistent with physical laws by treating physical evolution as a chain of causally connected events. It employs three modules: Physics-driven Event Chain Reasoning to decompose phenomena into scene-graph events, Transition-aware Routed Keyframe Conditioning to guide keyframe synthesis for smooth transitions, and Physics-injected Contrastive Semantic Guidance to steer generation toward plausible dynamics. Experiments on multiple physics benchmarks show improved physical plausibility compared to prior approaches.
By Zixuan Wang, Yixin Hu, Wen Li, Feng Chen, Yan Liu, Duo Peng, Yinjie Lei
WorldSculpt presents a method for generating compositional 3D representations of cluttered scenes with hundreds of objects by adapting a single-object 3D generative prior to multi-view observations. The approach, built on Pixal3D with a multi-view conditioning pathway, can generalize to highly occluded scenes without scene-level training. The authors also introduce the UE-MeshyScene benchmark and demonstrate that their method outperforms prior approaches across various evaluation settings, including converting existing 3DGS worlds into compositional mesh scenes.
By Muyao Niu, Jixuan He, Ruihan Yu, Lian Fu, Yonghao Yu, Zheng-Hui Huang, Yifan Zhan, Fengbo Lan, Yongtao Ge, Yinqiang Zheng, Kaipeng Zhang, Zhixiang Wang