Hugging Face Trending Papers
Jul 2

SimWorlds: A Multi-Agent System for Dynamic 3D Scene Creation

LLM agents are increasingly used to translate natural language into 3D scenes in a procedural way, but existing systems focus on static output. Dynamic 4D scenes from text alone, in which liquids flow, particles emit, rigid bodies cascade, and articulated mechanisms move, remain largely unexplored despite their value as editable content and as physics-grounded training data for video generation and embodied AI.

arXiv AI
Jun 2

Coding Agent Is Good As World Simulator

arXiv:2605. 14398v2 Announce Type: replace Abstract: World models have emerged as a powerful paradigm for building interactive simulation environments, with recent video-based approaches demonstrating impressive progress in generating visually plausible dynamics.

By Hongyu Wang, Jingquan Wang, Bocheng Zou, Radu Serban, Dan Negrut
arXiv Computation and Language
Sep 4

Editable Visual Design

Editable Visual Design introduces a new design paradigm that combines a Coding Agent with a Vision‑Language Model (VLM) and an image generation model. The VLM acts as the creative brain, understanding requirements, planning tasks, and judging aesthetics, while the image generator produces isolated visual assets on demand. The agent follows an "imagine first, then act" workflow, generating assets, writing native HTML/CSS, and refining the design through visual feedback, ultimately producing editable, layer‑wise artifacts with real text that can be adjusted via a graphical interface.

By Junyan Ye, Wei Liu, Dongzhi Jiang, Zichen Wen, HaoDong Li, Zhutao Lv, Jiaxin Lin, Jinhua Yu, Jun He, Zilong Huang, Rui Chen, Weijia Li
Hugging Face Trending Papers
Jun 25

In-Context Model Predictive Generation: Open-Vocabulary Motion Synthesis from Language Models to Physics

Synthesizing human motion from textual descriptions is essential for immersive digital applications, yet existing methods face a persistent trade-off between semantic fidelity and physical realism. Large language model (LLM)-based approaches can interpret diverse open-vocabulary instructions and compose high-level action plans, but they often generate motions that violate physical constraints.

arXiv AI
Aug 21

ChronoAgentic: A Code-based Multi-Agent World Simulator for Physically Grounded Simulation Construction

arXiv:2605. 14398v3 Announce Type: replace Abstract: Video-based world models generate visually plausible rollouts, but since they infer dynamics in latent states, they enforce no explicit physical constraints: contacts drift, shapes distort, and motion loses consistency.

By Hongyu Wang, Jingquan Wang, Ashvin Anilkumar, Bocheng Zou, Radu Serban, Dan Negrut