Programmable World Model
arXiv:2609.10540v1 Announce Type: new Abstract: Recent video world models generate increasingly realistic and interactive visual experiences, yet lack reliable mechanisms for maintaining persistent w...
arXiv:2606. 01869v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly asked not only to write static interfaces, but to construct executable interactive worlds from natural language.
arXiv:2609.10540v1 Announce Type: new Abstract: Recent video world models generate increasingly realistic and interactive visual experiences, yet lack reliable mechanisms for maintaining persistent w...
arXiv:2606. 01057v1 Announce Type: cross Abstract: Procedural 3D modeling through code is emerging as a versatile paradigm, offering deterministic, engine-ready, and precisely editable assets that neural 3D generators inherently lack.
RILA is an execution‑driven agent that integrates browser rendering into the generation loop for interactive web development. It uses an Action Interaction Verification module to replay reference interactions on generated pages, collecting execution‑aware observations, and an Execution‑aware Rendering Score to jointly assess interaction correctness and visual fidelity during iterative optimization. A data synthesis pipeline further augments training data, enabling RILA to significantly improve interaction and visual quality across foundation models, even outperforming larger one‑shot generators.
arXiv:2605. 14398v3 Announce Type: replace Abstract: Video-based world models generate visually plausible rollouts, but since they infer dynamics in latent states, they enforce no explicit physical constraints: contacts drift, shapes distort, and motion loses consistency.
arXiv:2609.21293v1 Announce Type: new Abstract: Autonomous software generation (ASG) aims to turn human requirements into executable applications, but delivering these applications does not necessari...
arXiv:2604. 14262v2 Announce Type: replace-cross Abstract: GUI grounding models report over 85% accuracy on standard benchmarks, yet drop 27-56 percentage points when instructions require spatial reasoning rather than direct element naming.
Code4Scene is a benchmark that evaluates coding agents on constructing and editing 3D scenes in Unreal Engine. It tests agents on two tasks: construction, where they must build a scene from open‑ended language, and editing, where they must recover a target scene from reference images while preserving everything else. The benchmark measures task fulfillment, artifact integrity, and physical validity, revealing that construction and editing performance are correlated but not interchangeable, with agents struggling most with spatial composition and precise edits.
arXiv:2607. 28645v1 Announce Type: cross Abstract: Recent multimodal large language models can convert visual designs directly into executable code, but real mobile products require multiple screenshots to become a buildable codebase with shared components and working navigation.
arXiv:2602. 18548v2 Announce Type: replace-cross Abstract: Design-to-code translates high-fidelity UI designs into executable front-end implementations, but progress remains hard to compare due to inconsistent datasets, toolchains, and evaluation protocols.
arXiv:2507.18625v3 Announce Type: replace-cross Abstract: Graphical user interface (UI) software has undergone a fundamental transformation from traditional two-dimensional (2D) desktop/web/mobile in...
arXiv:2603. 03482v2 Announce Type: replace-cross Abstract: Interactive world models continually generate video by responding to a user's actions, enabling open-ended generation capabilities.
arXiv:2609.22000v1 Announce Type: new Abstract: Computer-use agents (CUAs) have advanced along two separate lines: graphical interaction and software development through code and the command line. Re...