arXiv Computer Vision

4DSynth: Controllable Procedural World Synthesis for Dynamic Embodied Simulation

4DSynth is a controllable procedural system that transforms natural-language descriptions, blueprint masks, or single photographs into editable 4D environments featuring explicit geometry, animated actors, collision-free trajectories, and physics-ready simulation states. The system unifies animation, camera planning, rendering, and task generation within a single geometry-grounded representation, enabling scalable creation of dynamic embodied simulation scenes. Using 4DSynth, the authors built 4DSynth-Nav, an interactive navigation benchmark that demonstrates the reproducibility and tunability of procedural failures across vision‑language models.

Hugging Face Trending Papers
Aug 27

4DSynth: Controllable Procedural World Synthesis for Dynamic Embodied Simulation

4DSynth is a controllable procedural system that transforms natural-language descriptions, blueprint masks, or single photographs into editable 4D environments featuring explicit geometry, animated actors, collision-free trajectories, and physics-ready simulation states. The system unifies animation, camera planning, rendering, and task generation within a single geometry-grounded representation, enabling scalable creation of diverse, interactive scenes. Using 4DSynth, the authors built 4DSynth-Nav, an interactive navigation benchmark that demonstrates the reproducibility of failures and tunable difficulty across three tiers for vision‑language models.

Hugging Face Trending Papers
Jul 2

SimWorlds: A Multi-Agent System for Dynamic 3D Scene Creation

LLM agents are increasingly used to translate natural language into 3D scenes in a procedural way, but existing systems focus on static output. Dynamic 4D scenes from text alone, in which liquids flow, particles emit, rigid bodies cascade, and articulated mechanisms move, remain largely unexplored despite their value as editable content and as physics-grounded training data for video generation and embodied AI.

arXiv Computer Vision
3d ago

CL4D: Contrastive Language-4D Pretraining for Vision-Language Reasoning in Dynamic Scenes

arXiv:2608.18734v2 Announce Type: replace Abstract: 4D understanding and reasoning is a fundamental capability for embodied AI agents operating in dynamic physical environments. However, existing vis...

By Kumal Hewagamage, Isuranga Senavirathne, Sasika Amarasinghe, Hasitha Gallella, Dulanga Weerakoon, Vigneshwaran Subbaraju, Ranga Rodrigo
arXiv AI
Jul 14

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

arXiv:2607. 11643v1 Announce Type: cross Abstract: Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints.

By Xinghang Li, Jun Guo, Qiwei Li, Long Qian, Hang Lai, Yueze Wang, Hongyu Yan, Jiahang Cao, Xi Chen, Jingen Qu, Jiaxi Song, Nan Sun, Hanye Zhao, Futeng Liu, Wanli Peng, Heyun Wang, Yunhong Wang, Caoyu Xia, Jack Zhao, Diyun Xiang, Hangjun Ye, Heng Qu, Huaping Liu, Jason Li
Hugging Face Trending Papers
Jul 7

SPEAR: A Simulator for Photorealistic Embodied AI Research

Interactive simulators have become powerful tools for training embodied agents and generating synthetic visual data, but existing photorealistic simulators suffer from limited generality, programmability, and rendering speed. We address these limitations by introducing SPEAR: A Simulator for Photorealistic Embodied AI Research.

arXiv AI
Aug 7

iARCS: Iterative Agentic RL for Controllable 3D Scene Generation

arXiv:2608. 06161v1 Announce Type: new Abstract: Synthetic 3D scene generation is increasingly used as a data source for computer vision and embodied AI, but existing generators often optimize perceptual realism without reliably satisfying task-critical functional constraints.

By Saugat Adhikari, Ashok Prasad Neupane, Pramish Paudel, Ajad Chhatkuli, Danda Pani Paudel
Hugging Face Trending Papers
Jun 25

In-Context Model Predictive Generation: Open-Vocabulary Motion Synthesis from Language Models to Physics

Synthesizing human motion from textual descriptions is essential for immersive digital applications, yet existing methods face a persistent trade-off between semantic fidelity and physical realism. Large language model (LLM)-based approaches can interpret diverse open-vocabulary instructions and compose high-level action plans, but they often generate motions that violate physical constraints.