arXiv Machine Learning

Design-MLLM: A Reinforcement Alignment Framework for Verifiable and Aesthetic Interior Design

arXiv:2603. 13312v2 Announce Type: replace-cross Abstract: Interior design is a requirements-to-visual-plan generation process that must simultaneously satisfy verifiable spatial feasibility and comparative aesthetic preferences.

arXiv AI
Aug 25

TO-Agents: A Multi-Agent AI Framework for Subjective Preference-Guided Topology Optimization

TO-Agents is a multi‑agent AI framework that translates natural‑language design intent into iterative topology optimization. It converts a human problem description into solver inputs, runs the optimizer, renders 3D topologies, and employs a judge agent to critique and revise results using multiview vision‑language reasoning. Evaluated on a cantilever beam and a phone‑stand design, the system achieved preference‑aligned designs in 60% of trials, outperforming an ablated pipeline by up to six times and enabling end‑to‑end intent‑to‑prototype design with additive manufacturing.

By Isabella A. Stewart, Hongrui Chen, Faez Ahmed
arXiv AI
Sep 10

Inferring the Unspoken: Aligning Embodied Agents with Implicit Preferences

The paper "Inferring the Unspoken: Aligning Embodied Agents with Implicit Preferences" addresses the challenge of natural-language instructions that omit details needed for embodied action. It introduces the Preference-based Planning (PbP) benchmark, comprising 5,000 evaluation groups and 290 preferences across three levels, to systematically evaluate agents’ ability to infer latent user preferences from a few demonstrations. The authors propose the two-stage Inferring the Unspoken (InTU) framework, which first verbalizes inferred preferences from multimodal demonstrations and then generates action plans conditioned on that explicit representation, showing that explicit verbalization improves alignment and robustness compared to direct end-to-end planning.

By Manjie Xu, Xinyi Yang, Wei Liang, Chi Zhang, Yixin Zhu
arXiv AI
Sep 15

Enabling Creative Exploration for Vibe Design Agents

The paper introduces a new inference architecture for vibe design agents that separates design exploration from implementation. By generating structured design specifications with typicality scores and selecting one for downstream generation, the method allows users to explore coherent UI alternatives without altering the underlying generation settings. Experiments on UI themes and visual-asset prompts show increased selection coverage and screenshot variation, with mixed preferences from an LLM judge and modest operational costs in a large online test.

By Yifan Zhang, Nghi D. Q. Bui, Georgios Evangelopoulos, Arnaud Benard
arXiv Computation and Language
Sep 1

PlanCraft: Sketch, Refine, and Furnish for Architect-Inspired Progressive 3D Residential Scene Generation

PlanCraft introduces a progressive approach to 3D residential scene generation that mirrors how architects design: starting with rough sketches and refining them over time. It leverages a large dataset of real floor plans to train a SketchPlan module that generates partial sketches at various completion levels, a PlanCraft‑Diff module that sharpens these sketches into precise vector floor plans, and a PlanCraft‑Agent that furnishes rooms within the established spatial contract. The method outperforms existing 2D and 3D baselines, achieving a 61.1% lower FID and a 15‑point lead in expert‑rated spatial rationality, even with only 25% sketch completion.

By Pengyu Zeng, Yuqin Dai, Jun Yin, Ziyang Han, Ng Cheuk Hei, Jing Zhong, Chaoyang Shi, ZhanXiang Jin, Maowei Jiang, Shuai Lu
arXiv Machine Learning
Jun 26

Staying VIGILant: Mitigating Visual Laziness via Counterfactual Visual Alignment in MLLMs

arXiv:2606. 26387v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) extend large language models (LLMs) with visual perception, enabling joint reasoning over images and text.

By Xi Xiao, Chen Liu, Chih-Ting Liao, Yunbei Zhang, Qizhen Lan, Yuxiang Wei, Lin Zhao, Janet Wang, Jianyang Gu, Muchao Ye, Tianyang Wang, Hao Xu