arXiv:2607. 22928v1 Announce Type: new Abstract: Generative UI tools promise to democratize UI design by turning natural language descriptions into complete interfaces.
By Kashif Imteyaz, Kaif Imteyaz, Nakul Rajpal, Kaif Shaikh, Michael Muller, Saiph Savage
arXiv:2606. 31711v1 Announce Type: new Abstract: Faithfulness -- how precisely a generated image aligns with its prompt -- is increasingly central to the real-world utility of text-to-image (T2I) models.
By Yuanhao Ban, Tong Xie, Sohyun An, Yunqi Hong, Evan Frick, I-Hung Hsu, Wei-Lin Chiang, Ion Stoica, Cho-Jui Hsieh
TO-Agents is a multi‑agent AI framework that translates natural‑language design intent into iterative topology optimization. It converts a human problem description into solver inputs, runs the optimizer, renders 3D topologies, and employs a judge agent to critique and revise results using multiview vision‑language reasoning. Evaluated on a cantilever beam and a phone‑stand design, the system achieved preference‑aligned designs in 60% of trials, outperforming an ablated pipeline by up to six times and enabling end‑to‑end intent‑to‑prototype design with additive manufacturing.
By Isabella A. Stewart, Hongrui Chen, Faez Ahmed
The paper "Inferring the Unspoken: Aligning Embodied Agents with Implicit Preferences" addresses the challenge of natural-language instructions that omit details needed for embodied action. It introduces the Preference-based Planning (PbP) benchmark, comprising 5,000 evaluation groups and 290 preferences across three levels, to systematically evaluate agents’ ability to infer latent user preferences from a few demonstrations. The authors propose the two-stage Inferring the Unspoken (InTU) framework, which first verbalizes inferred preferences from multimodal demonstrations and then generates action plans conditioned on that explicit representation, showing that explicit verbalization improves alignment and robustness compared to direct end-to-end planning.
By Manjie Xu, Xinyi Yang, Wei Liang, Chi Zhang, Yixin Zhu
arXiv:2607. 05573v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) and Vision-Language Models (VLMs) enable the automatic generation of parametric 3D designs from natural-language specifications.
By J de Curt\`o, Victoria Guill\'en, I. de Zarz\`a
AI agents that generate final answers based on user input often do not meet the needs of creative fields. Fields such as structural design and architecture need interactive systems that help users externalise and develop ideas, explore alternatives, and refine partial solutions.