Design Theater: A Benchmark for Generative UI
arXiv:2607. 22928v1 Announce Type: new Abstract: Generative UI tools promise to democratize UI design by turning natural language descriptions into complete interfaces.
arXiv:2603. 13312v2 Announce Type: replace-cross Abstract: Interior design is a requirements-to-visual-plan generation process that must simultaneously satisfy verifiable spatial feasibility and comparative aesthetic preferences.
arXiv:2607. 22928v1 Announce Type: new Abstract: Generative UI tools promise to democratize UI design by turning natural language descriptions into complete interfaces.
arXiv:2606. 31711v1 Announce Type: new Abstract: Faithfulness -- how precisely a generated image aligns with its prompt -- is increasingly central to the real-world utility of text-to-image (T2I) models.
TO-Agents is a multi‑agent AI framework that translates natural‑language design intent into iterative topology optimization. It converts a human problem description into solver inputs, runs the optimizer, renders 3D topologies, and employs a judge agent to critique and revise results using multiview vision‑language reasoning. Evaluated on a cantilever beam and a phone‑stand design, the system achieved preference‑aligned designs in 60% of trials, outperforming an ablated pipeline by up to six times and enabling end‑to‑end intent‑to‑prototype design with additive manufacturing.
The paper "Inferring the Unspoken: Aligning Embodied Agents with Implicit Preferences" addresses the challenge of natural-language instructions that omit details needed for embodied action. It introduces the Preference-based Planning (PbP) benchmark, comprising 5,000 evaluation groups and 290 preferences across three levels, to systematically evaluate agents’ ability to infer latent user preferences from a few demonstrations. The authors propose the two-stage Inferring the Unspoken (InTU) framework, which first verbalizes inferred preferences from multimodal demonstrations and then generates action plans conditioned on that explicit representation, showing that explicit verbalization improves alignment and robustness compared to direct end-to-end planning.
arXiv:2607. 05573v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) and Vision-Language Models (VLMs) enable the automatic generation of parametric 3D designs from natural-language specifications.
AI agents that generate final answers based on user input often do not meet the needs of creative fields. Fields such as structural design and architecture need interactive systems that help users externalise and develop ideas, explore alternatives, and refine partial solutions.
arXiv:2505. 05026v5 Announce Type: replace-cross Abstract: User interface (UI) design goes beyond visuals to shape user experience (UX), underscoring the shift toward UI/UX as a unified concept.
arXiv:2607. 07521v1 Announce Type: cross Abstract: AI agents that generate final answers based on user input often do not meet the needs of creative fields.
The paper introduces a new inference architecture for vibe design agents that separates design exploration from implementation. By generating structured design specifications with typicality scores and selecting one for downstream generation, the method allows users to explore coherent UI alternatives without altering the underlying generation settings. Experiments on UI themes and visual-asset prompts show increased selection coverage and screenshot variation, with mixed preferences from an LLM judge and modest operational costs in a large online test.
PlanCraft introduces a progressive approach to 3D residential scene generation that mirrors how architects design: starting with rough sketches and refining them over time. It leverages a large dataset of real floor plans to train a SketchPlan module that generates partial sketches at various completion levels, a PlanCraft‑Diff module that sharpens these sketches into precise vector floor plans, and a PlanCraft‑Agent that furnishes rooms within the established spatial contract. The method outperforms existing 2D and 3D baselines, achieving a 61.1% lower FID and a 15‑point lead in expert‑rated spatial rationality, even with only 25% sketch completion.
arXiv:2506.02015v4 Announce Type: replace Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have enabled unified multimodal understanding and generation. However, they still strug...
arXiv:2606. 26387v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) extend large language models (LLMs) with visual perception, enabling joint reasoning over images and text.