arXiv AI

Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes

Code4Scene is a benchmark that evaluates coding agents on constructing and editing 3D scenes in Unreal Engine. It tests agents on two tasks: construction, where they must build a scene from open‑ended language, and editing, where they must recover a target scene from reference images while preserving everything else. The benchmark measures task fulfillment, artifact integrity, and physical validity, revealing that construction and editing performance are correlated but not interchangeable, with agents struggling most with spatial composition and precise edits.

arXiv Computer Vision
Aug 28

Procedura: Agentic 3D Modeling with Procedural Control

Procedura is a new 3D modeling agent that treats 3D shape as code, using a large language model to generate a procedural assembly from a text prompt. It constructs an assembly graph, writes a parametric program with named parts and typed mates, and verifies each part through compile, mate, and connectivity checks before adding it. A vision critic refines the assembly step‑by‑step, and the resulting program includes per‑part materials and simulator‑validated articulation, producing sharp edges and editable, part‑structured outputs that outperform existing native 3D generators on P3D‑Bench and MechBench‑36.

By Youtian Lin, Yikang Yang, Zhanpeng Hu, Mengqi Zhou, Feihu Zhang, Xun Cao, Jiaheng Liu, Yao Yao
arXiv AI
Aug 21

ChronoAgentic: A Code-based Multi-Agent World Simulator for Physically Grounded Simulation Construction

arXiv:2605. 14398v3 Announce Type: replace Abstract: Video-based world models generate visually plausible rollouts, but since they infer dynamics in latent states, they enforce no explicit physical constraints: contacts drift, shapes distort, and motion loses consistency.

By Hongyu Wang, Jingquan Wang, Ashvin Anilkumar, Bocheng Zou, Radu Serban, Dan Negrut
arXiv AI
Sep 25

PPTBench: Can Coding Agents Reconstruct the Visual World through Structured, Editable Slides

PPTBench is a new benchmark that tests coding agents’ ability to reconstruct scientific flow diagrams from arXiv papers into editable PowerPoint slides. The dataset contains 500 tasks, each requiring agents to produce a single PPTX page with native, editable objects, and a four‑stage Agentic Judge evaluates validity, semantic correctness, rendering quality, and fine‑grained visual quality. Across 31 model configurations, the best score is 67.80, with a median of 19.47, showing that while agents can generate valid PPTX files, they still struggle with semantic and visual accuracy, especially text details.

By Xiaoqiu Wang, Yizhe Chi, Wenyi Li, Deyao Hong, Zhihan Shan, Mingju Gao, Kaisen Yang, Youjie Zheng, Calvin Xiao, Qinhuai Na
arXiv AI
Aug 11

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

arXiv:2607. 27380v2 Announce Type: replace-cross Abstract: Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt.

By Haodong Li, Tianfei Ren, Xiaoxiao Ma, Chunmei Qing, Zhen Fang, Sipeng He, Ziyu Guo, Haoyu Wu, Juanxi Tian, Yihang Zou, Ruichuan An, Dongzhi Jiang, Boxue Yang, Ji Xie, Xu Huang, Wenhao Yan, Jialv Zou, Zhengrong Yue, Yaxin Luo, Xiaotong Li, Yuzhu Wang, Junyan Ye, Jinjing Zhao, Zehui Chen, Lin Chen, Renye Yan, Feng Zhao, Pheng-Ann Heng
arXiv Computation and Language
Sep 1

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence

The paper surveys Multimodal Code Intelligence, focusing on tasks where code is generated, edited, refined, or reasoned about under visually grounded inputs such as screenshots, charts, and videos. It categorizes the field by the role of code—rendered artifact, editable structure, intermediate reasoning trace, or executable tool interface—and organizes benchmarks into four domains: Graphical User Interface, Scientific Visualization, Structured Graphics, and Frontier Tasks and Frameworks. The authors argue that reliable evaluation must include evidence of semantics and interaction beyond visual fidelity, and propose four verification-centered research directions to advance the field toward evidence-grounded executable systems.

By Xuanle Zhao, Qiushi Sun, Jingyu Xiao, Xuexin Liu, Haoyue Yang, Qiaosheng Chen, Xianzhen Luo, Jing Huang, Yufeng Zhong, Lei Chen, Shuai Fu, Zhenlin Wei, Jinhe Bi, Lei Jiang, Haibo Qiu, Siqi Yang, Peng Shi, Jian Hu, Zhixiong Zeng
arXiv Computer Vision
Sep 25

AgenticCADedit: A Stateful, Tool-Mediated Agentic Approach to Multimodal 3D CAD Editing

AgenticCADedit introduces a stateful, tool‑mediated approach to multimodal 3D CAD editing, transforming the process from generating a single complete program to executing a sequence of incremental, verifiable actions on a persistent CAD state. By committing each step, inspecting geometry, and selectively reverting faulty operations, the method preserves partial progress and builds upon earlier edits. Experiments across three large language models show substantial gains in validity and acceptance, with the weakest baseline model’s validity rising from 51.0% to 94.8% and a token‑cost reduction of 66.7% compared to neuralCAD‑Edit.

By Saptarshi Neil Sinha, Mika Silvan Goschke, Paul Julius K\"uhn, Arjan Kuijper, Michael Weinmann
Hugging Face Trending Papers
3d ago

Timeline-Bench: Evaluating Agents on Realistic Video-Editing Tasks, from Raw Footage to Final Cut

Timeline-Bench is a benchmark comprising 56 real video‑editing tasks that require AI agents to transform raw production material into finished videos. Each task includes a brief, source assets, a container, and a set of tests that assess format, content, brief compliance, and quality based on 2,582 blind judgments by 43 video editors. In evaluations, the best agent resolved only 15 of the 56 tasks, and most failures were due to quality tests rather than technical errors.