EngiWorld is a new benchmark that tests autonomous agents across the full engineering design loop, covering 1,301 expert‑curated tasks in six domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms with both GUI and CLI interfaces. It introduces an artifact‑centric evaluation method that programmatically verifies geometric validity, physical feasibility, and rule compliance of both final and intermediate artifacts, scoring tasks continuously rather than with binary success. Initial tests of seven frontier models show a large capability gap, with the best model scoring only 44.3 on the EngiScore and just 3.6% of multi‑software attempts succeeding.
By Hongcheng Gao, Hailong Qu, Yu Lei, Henghui Sun, Haoyang Li, Yipeng Wei, Naihao Xue, Xiaohan Yu, Zhuo Tao, Yihe Zang, Yajiao Wang, Jingyi Tang, Yi Li, Jingjing Zhou, Jie Luo, Bohan Zeng, Chengyu Shen, Hao Jiang, Chong Chen, Bowen Qu, Olive Huang, Zeqiang Wang
arXiv:2607. 17528v1 Announce Type: new Abstract: LLM-driven agent systems have emerged as a promising paradigm for electronic design automation (EDA), demonstrating strong potential for automating complex design workflows.
By Jinyuan Deng, Zhengrui Chen, Xufeng Wei, Tianyu Xing, Chenyi Wen, Cheng Zhuo
arXiv:2609.33772v2 Announce Type: replace
Abstract: Executable environments are critical for post-training agents on tasks that require tool use and multi-step interaction, but constructing executabl...
By Weiyi Xu, Xiaowen Yang, Wen Da, Hang Xu, Canwei Li, Hongjie You, Pusen Dong, Yucheng Zeng, Zhaokai Luo, Mu Chuan
arXiv:2607. 14896v1 Announce Type: cross Abstract: Addressing a structural-engineering request requires more than a single answer; it requires a chain of interdependent artifacts: interpreted requirements, a computable model, validation records, solver outputs, code-check records, and a final report.
By Sizhong Qin, Yi Gu, Yao Jiang, Ao Cai, Changjian Zhou, Shaoxuan Shuai, Jiachang Wang, Tianhao Shen, Yueqiang Li, Xinhao Li, Li Zeng, Yueshi Chen, Dachen Gao, Genrong Xu, Wenjie Liao, Xinzheng Lu
arXiv:2606. 30573v1 Announce Type: new Abstract: We introduce SWE-Interact, a new testbed for evaluating coding agents on multi-turn, interactive, user-driven software engineering tasks.
By Mohit Raghavendra, Anisha Gunjal, Aakash Sabharwal, Yunzhong He
arXiv:2607. 23124v1 Announce Type: new Abstract: Large language model agents have advanced rapidly, yet progress remains fragmented across domains, capabilities, task difficulty, and interaction settings.
By Hao Jiang, Gangtao Xin, Yingdi Huang, Guojie Zhu, Jiangshan Zhang, Xinyuan Lin, Yunkun Xu, Chengyu Shen, Wenlong Fei, Jiawei Li, Yujie Fu, Sichen Kang, Tingyu Xie, Yedi Hu, Jingren Zhang, Hongcheng Gao, Jianshu Zeng, Chong Chen, Chang Guo, Chao Feng, Feng Wang, Fulin Lin, Jinchao Ma, Lang Mei, Li Huang, Liyan Liu, Qing He, Shuting Tao, Siyu Mo, Xiangnan Chen, Xiaohan Yu, Xiaoyang Li, Yanheng Hou, Yanyu Wu, Zhihan Yang, Wentao Zhang, Yang Gao, Zhao Cao