arXiv AI

Design Theater: A Benchmark for Generative UI

arXiv:2607. 22928v1 Announce Type: new Abstract: Generative UI tools promise to democratize UI design by turning natural language descriptions into complete interfaces.

arXiv AI
Sep 15

Enabling Creative Exploration for Vibe Design Agents

The paper introduces a new inference architecture for vibe design agents that separates design exploration from implementation. By generating structured design specifications with typicality scores and selecting one for downstream generation, the method allows users to explore coherent UI alternatives without altering the underlying generation settings. Experiments on UI themes and visual-asset prompts show increased selection coverage and screenshot variation, with mixed preferences from an LLM judge and modest operational costs in a large online test.

By Yifan Zhang, Nghi D. Q. Bui, Georgios Evangelopoulos, Arnaud Benard
arXiv AI
Sep 15

Efficient Personalization of Generative User Interfaces

The paper "Efficient Personalization of Generative User Interfaces" addresses the challenge of tailoring generative user interfaces (GenUIs) to individual users when interface screens are not pre‑defined. By collecting judgments from 20 participants on 600 GenUI pairs, the authors show low agreement (Krippendorff's alpha = 0.25) and diverse rationales for UI preferences. They propose a sample‑efficient personalization method that leverages a few pairwise judgments to weight prior users’ preferences, outperforming a pretrained UI evaluator and a larger multimodal model offline and outperforming all baselines in an online study with 12 new users. "whyItMatters":"The study demonstrates a practical approach to personalizing on‑demand interfaces, showing that even sparse, subjective feedback can be effectively used to improve user satisfaction with generative UI designs."

By Yi-Hao Peng, Jeffrey P. Bigham, Jason Wu
arXiv Computation and Language
Sep 4

Editable Visual Design

Editable Visual Design introduces a new design paradigm that combines a Coding Agent with a Vision‑Language Model (VLM) and an image generation model. The VLM acts as the creative brain, understanding requirements, planning tasks, and judging aesthetics, while the image generator produces isolated visual assets on demand. The agent follows an "imagine first, then act" workflow, generating assets, writing native HTML/CSS, and refining the design through visual feedback, ultimately producing editable, layer‑wise artifacts with real text that can be adjusted via a graphical interface.

By Junyan Ye, Wei Liu, Dongzhi Jiang, Zichen Wen, HaoDong Li, Zhutao Lv, Jiaxin Lin, Jinhua Yu, Jun He, Zilong Huang, Rui Chen, Weijia Li
arXiv AI
Aug 25

SlideGen: Collaborative Multimodal Agents for Scientific Slide Generation

SlideGen is a collaborative vision‑language multi‑agent framework designed to generate scientific presentation slides from research papers. It assigns specialized agents to outline the presentation structure, align figures and tables with key claims, generate speaker notes, and compose editable PPTX slides using a diverse layout library. The system introduces a geometry‑aware density metric to evaluate visual clutter and demonstrates significant improvements in layout balance, content coverage, and text coherence over existing baselines on a 200‑paper benchmark.

By Xin Liang, Zhilin Zhang, Xiang Zhang, Haoran Su, Yiwei Xu, Siqi Sun, Chenyu You