Minimal Local Simulation Foundations for LLM- and VLM-Driven Agents in 2D and 3D Environments
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The paper introduces two minimal simulation foundations—SD-AgentFoundry-2D and SD-AgentFoundry-3D—for educational and rapid prototyping use with large language models (LLMs) and vision-language models (VLMs). SD-AgentFoundry-2D offers a 2D multi‑agent environment where LLM agents move, communicate, and react to local events such as fire, while SD-AgentFoundry-3D provides a 3D digital‑twin setting where a VLM interprets first‑person images to generate natural‑language movement instructions. Both frameworks run locally on macOS, Windows, and Linux, are intentionally lightweight, and are open for modification rather than being finished applications.
arXiv:2607. 00989v1 Announce Type: cross Abstract: Semantic trajectory analysis has recently emerged as an approach for modeling human movement by capturing implicit patterns and behaviors through semantic information (e.
arXiv:2607. 21522v1 Announce Type: cross Abstract: Creating dynamic and physically realistic 4D worlds from natural language descriptions is both fascinating and challenging.
arXiv:2507. 09788v3 Announce Type: replace-cross Abstract: Recent advances in Large Language Models (LLM) have led to a new class of autonomous agents, renewing and expanding interest in the area.
LLM agents are increasingly used to translate natural language into 3D scenes in a procedural way, but existing systems focus on static output. Dynamic 4D scenes from text alone, in which liquids flow, particles emit, rigid bodies cascade, and articulated mechanisms move, remain largely unexplored despite their value as editable content and as physics-grounded training data for video generation and embodied AI.
arXiv:2605. 14398v2 Announce Type: replace Abstract: World models have emerged as a powerful paradigm for building interactive simulation environments, with recent video-based approaches demonstrating impressive progress in generating visually plausible dynamics.