arXiv AI

EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling

EditHero is presented as the first benchmark for long-horizon, part-level 3D editing, featuring natural-language instructions and target images for both geometry and texture. The benchmark uses a deterministic assembly engine that produces the exact target after each edit, and every sequence is manually reviewed. It compares non-agentic top‑down methods with LLM/VLM agent bottom‑up approaches, finding that the latter follow instructions more closely and preserve unedited parts better, though each edit takes minutes.

arXiv Computer Vision
Aug 28

Procedura: Agentic 3D Modeling with Procedural Control

Procedura is a new 3D modeling agent that treats 3D shape as code, using a large language model to generate a procedural assembly from a text prompt. It constructs an assembly graph, writes a parametric program with named parts and typed mates, and verifies each part through compile, mate, and connectivity checks before adding it. A vision critic refines the assembly step‑by‑step, and the resulting program includes per‑part materials and simulator‑validated articulation, producing sharp edges and editable, part‑structured outputs that outperform existing native 3D generators on P3D‑Bench and MechBench‑36.

By Youtian Lin, Yikang Yang, Zhanpeng Hu, Mengqi Zhou, Feihu Zhang, Xun Cao, Jiaheng Liu, Yao Yao
arXiv Computer Vision
Sep 25

AgenticCADedit: A Stateful, Tool-Mediated Agentic Approach to Multimodal 3D CAD Editing

AgenticCADedit introduces a stateful, tool‑mediated approach to multimodal 3D CAD editing, transforming the process from generating a single complete program to executing a sequence of incremental, verifiable actions on a persistent CAD state. By committing each step, inspecting geometry, and selectively reverting faulty operations, the method preserves partial progress and builds upon earlier edits. Experiments across three large language models show substantial gains in validity and acceptance, with the weakest baseline model’s validity rising from 51.0% to 94.8% and a token‑cost reduction of 66.7% compared to neuralCAD‑Edit.

By Saptarshi Neil Sinha, Mika Silvan Goschke, Paul Julius K\"uhn, Arjan Kuijper, Michael Weinmann
arXiv Computer Vision
4d ago

EditVerse3D: High-Quality 3D Object Editing with Region-Aware Learning

arXiv:2607.07187v2 Announce Type: replace Abstract: Local editing of 3D objects remains a long-standing challenge. When interacting with 3D content, humans naturally tend to specify a coarse region o...

By Youtan Yin, Yanning Zhou, Jiacheng Wei, Xiaofeng Yang, Jun Zhang, Jiayang Bai, Jingwen Ye, Weidong Zhang, Guosheng Lin
arXiv AI
6d ago

Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes

Code4Scene is a benchmark that evaluates coding agents on constructing and editing 3D scenes in Unreal Engine. It tests agents on two tasks: construction, where they must build a scene from open‑ended language, and editing, where they must recover a target scene from reference images while preserving everything else. The benchmark measures task fulfillment, artifact integrity, and physical validity, revealing that construction and editing performance are correlated but not interchangeable, with agents struggling most with spatial composition and precise edits.

By Xiaokang Ye, Siddhant Hitesh Mantri, Zimeng Chen, Edward Zhang, Zhaoxu Zheng, Yuanheng Li, Yizhao Chen, Tianyang Huang, Lianhui Qin