arXiv:2607. 03006v1 Announce Type: cross Abstract: Text-rich image models can now design poster-scale layouts, but we lack ways to measure whether they honor scientific communication contracts: legible labels, prescribed aspect ratios, and -- above all -- abstaining from fabricated scientific figures.
By Tianyi Yang, Dawei Fu, Youpeng Wu, Zixun Kou, Linrui Chen, Ruobing Jiang, Zijian Wang, Qiang Li
AgenticCADedit introduces a stateful, tool‑mediated approach to multimodal 3D CAD editing, transforming the process from generating a single complete program to executing a sequence of incremental, verifiable actions on a persistent CAD state. By committing each step, inspecting geometry, and selectively reverting faulty operations, the method preserves partial progress and builds upon earlier edits. Experiments across three large language models show substantial gains in validity and acceptance, with the weakest baseline model’s validity rising from 51.0% to 94.8% and a token‑cost reduction of 66.7% compared to neuralCAD‑Edit.
By Saptarshi Neil Sinha, Mika Silvan Goschke, Paul Julius K\"uhn, Arjan Kuijper, Michael Weinmann
ACE is a self‑correcting agentic canvas editor that operates on a hierarchical scene‑graph rather than flat document formats, enabling reliable multi‑slide presentation automation. It pairs a presentation‑specialized action space of 98 tools with CARE, a content‑aware router that reduces input tokens by about 89%, and a ground‑truth‑free instruction‑following judge that feeds natural‑language critiques back into the agent for self‑correction. In benchmarks, ACE outperforms a comparable agentic HTML pipeline on instruction following (4.23 vs. 3.81), runs 1.75× faster, costs 44% less, and is preferred by 58.7% of blind raters, with 81% favoring the self‑corrected output.
By JooYoung Jang, Taegyeong Lee, Jihyeon Park, Nojun Kwak
Professional graphic design is a long-horizon agentic task in which structured, editable artifacts emerge from many interdependent actions, yet outcomes admit no reliable programmatic oracle. We intro...
BlueprintAgent (BPA) is a multimodal agent that converts scanned reinforced‑concrete building blueprints into simulation‑ready frame models by treating a large language model (MLLM) as the primary reader and decision maker, while OCR and computer vision provide localized evidence. BPA implements engineering constraints as callable validators that trigger targeted MLLM revisits over specific regions, enabling precise beam‑column support, span count, and 3D continuity checks. In evaluations on 300 real scanned sheets from 20 projects, BPA achieved a macro‑averaged Beam F1 of 0.994, vastly outperforming single‑MLLM zero‑shot (0.301) and fixed‑pipeline (0.820) baselines.
By Zhouyuan Xu, Chen Yang, Linhao Wang, Jiansheng Fan, Chen Wang
Designer‑RSI presents a continual adaptation framework that lets a frozen frontier model operate professional design software while an external procedural memory learns natural‑language design skills from user traffic. Over five rounds on 1,406 real briefs and 1,869 graded trajectories, the memory grew from 76 to 139 skills, boosting execution success from 72.7% to 99.3% and improving win rates on four design benchmarks. The study shows that widening and deepening the memory, especially together, significantly outperforms a no‑skill baseline.
By Hongyang Du, Lan Yan, Christian Flores, Asim Kadav