ESG: Generating Physically Consistent Dynamic 3D Scenes from Text Descriptions
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
LLM agents are increasingly used to translate natural language into 3D scenes in a procedural way, but existing systems focus on static output. Dynamic 4D scenes from text alone, in which liquids flow, particles emit, rigid bodies cascade, and articulated mechanisms move, remain largely unexplored despite their value as editable content and as physics-grounded training data for video generation and embodied AI.
arXiv:2607. 01766v1 Announce Type: new Abstract: LLM agents are increasingly used to translate natural language into 3D scenes in a procedural way, but existing systems focus on static output.
The paper introduces LEGO, a benchmark dataset pairing user text descriptions with human‑annotated fine‑grained constraints and reference 3D scenes, and LEGO‑Eval, an evaluation framework that decomposes descriptions into atomic constraints and verifies each using grounding and spatial reasoning tools. It demonstrates that LEGO‑Eval detects misalignment more accurately than existing methods and that current 3D scene synthesis approaches achieve at most a 10% success rate on this benchmark.
arXiv:2607. 27380v2 Announce Type: replace-cross Abstract: Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt.
arXiv:2602. 10840v2 Announce Type: replace Abstract: Large language models (LLMs) have been widely studied in areas such as mathematical reasoning, complex coding, and scientific problem solving.
4DSynth is a controllable procedural system that transforms natural-language descriptions, blueprint masks, or single photographs into editable 4D environments featuring explicit geometry, animated actors, collision-free trajectories, and physics-ready simulation states. The system unifies animation, camera planning, rendering, and task generation within a single geometry-grounded representation, enabling scalable creation of dynamic embodied simulation scenes. Using 4DSynth, the authors built 4DSynth-Nav, an interactive navigation benchmark that demonstrates the reproducibility and tunability of procedural failures across vision‑language models.