arXiv AI By Gioele Molinari, Florian Felten, Soheyl Massoudi, Mark Fuge

EngiAI: Capability-Based Evaluation of Tool-Connected LLM Agents for Engineering Design

Read the original on arXiv AI →

EngiAI introduces a capability-based evaluation framework for tool-connected engineering agents, assessing workflow execution, retrieval-assisted parameter selection, HPC orchestration, and training-code authoring using execution traces and engineering artifacts. The framework was applied to four LLM backends on EngiBench Beams2D and Photonics2D, revealing that proprietary models outperform open-source ones in workflow completion and HPC orchestration, while indexed retrieval improves parameter selection. The study demonstrates that evaluating distinct skills separately provides clearer insight into failure mechanisms than end-to-end success rates alone.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments?

EngiWorld is a new benchmark that tests autonomous agents across the full engineering design loop, covering 1,301 expert‑curated tasks in six domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms with both GUI and CLI interfaces. It introduces an artifact‑centric evaluation method that programmatically verifies geometric validity, physical feasibility, and rule compliance of both final and intermediate artifacts, scoring tasks continuously rather than with binary success. Initial tests of seven frontier models show a large capability gap, with the best model scoring only 44.3 on the EngiScore and just 3.6% of multi‑software attempts succeeding.

By Hongcheng Gao, Hailong Qu, Yu Lei, Henghui Sun, Haoyang Li, Yipeng Wei, Naihao Xue, Xiaohan Yu, Zhuo Tao, Yihe Zang, Yajiao Wang, Jingyi Tang, Yi Li, Jingjing Zhou, Jie Luo, Bohan Zeng, Chengyu Shen, Hao Jiang, Chong Chen, Bowen Qu, Olive Huang, Zeqiang Wang
arXiv AI
Jul 17

StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows

arXiv:2607. 14896v1 Announce Type: cross Abstract: Addressing a structural-engineering request requires more than a single answer; it requires a chain of interdependent artifacts: interpreted requirements, a computable model, validation records, solver outputs, code-check records, and a final report.

By Sizhong Qin, Yi Gu, Yao Jiang, Ao Cai, Changjian Zhou, Shaoxuan Shuai, Jiachang Wang, Tianhao Shen, Yueqiang Li, Xinhao Li, Li Zeng, Yueshi Chen, Dachen Gao, Genrong Xu, Wenjie Liao, Xinzheng Lu
arXiv AI
Jul 28

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications

arXiv:2607. 23124v1 Announce Type: new Abstract: Large language model agents have advanced rapidly, yet progress remains fragmented across domains, capabilities, task difficulty, and interaction settings.

By Hao Jiang, Gangtao Xin, Yingdi Huang, Guojie Zhu, Jiangshan Zhang, Xinyuan Lin, Yunkun Xu, Chengyu Shen, Wenlong Fei, Jiawei Li, Yujie Fu, Sichen Kang, Tingyu Xie, Yedi Hu, Jingren Zhang, Hongcheng Gao, Jianshu Zeng, Chong Chen, Chang Guo, Chao Feng, Feng Wang, Fulin Lin, Jinchao Ma, Lang Mei, Li Huang, Liyan Liu, Qing He, Shuting Tao, Siyu Mo, Xiangnan Chen, Xiaohan Yu, Xiaoyang Li, Yanheng Hou, Yanyu Wu, Zhihan Yang, Wentao Zhang, Yang Gao, Zhao Cao