Minimal Local Simulation Foundations for LLM- and VLM-Driven Agents in 2D and 3D Environments
Read the original on arXiv AI →The paper introduces two minimal simulation foundations—SD-AgentFoundry-2D and SD-AgentFoundry-3D—for educational and rapid prototyping use with large language models (LLMs) and vision-language models (VLMs). SD-AgentFoundry-2D offers a 2D multi‑agent environment where LLM agents move, communicate, and react to local events such as fire, while SD-AgentFoundry-3D provides a 3D digital‑twin setting where a VLM interprets first‑person images to generate natural‑language movement instructions. Both frameworks run locally on macOS, Windows, and Linux, are intentionally lightweight, and are open for modification rather than being finished applications.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.