arXiv AI

EngiAI: Capability-Based Evaluation of Tool-Connected LLM Agents for Engineering Design

EngiAI introduces a capability-based evaluation framework for tool-connected engineering agents, assessing workflow execution, retrieval-assisted parameter selection, HPC orchestration, and training-code authoring using execution traces and engineering artifacts. The framework was applied to four LLM backends on EngiBench Beams2D and Photonics2D, revealing that proprietary models outperform open-source ones in workflow completion and HPC orchestration, while indexed retrieval improves parameter selection. The study demonstrates that evaluating distinct skills separately provides clearer insight into failure mechanisms than end-to-end success rates alone.

arXiv AI
2d ago

EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments?

EngiWorld is a new benchmark that tests autonomous agents across the full engineering design loop, covering 1,301 expert‑curated tasks in six domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms with both GUI and CLI interfaces. It introduces an artifact‑centric evaluation method that programmatically verifies geometric validity, physical feasibility, and rule compliance of both final and intermediate artifacts, scoring tasks continuously rather than with binary success. Initial tests of seven frontier models show a large capability gap, with the best model scoring only 44.3 on the EngiScore and just 3.6% of multi‑software attempts succeeding.

By Hongcheng Gao, Hailong Qu, Yu Lei, Henghui Sun, Haoyang Li, Yipeng Wei, Naihao Xue, Xiaohan Yu, Zhuo Tao, Yihe Zang, Yajiao Wang, Jingyi Tang, Yi Li, Jingjing Zhou, Jie Luo, Bohan Zeng, Chengyu Shen, Hao Jiang, Chong Chen, Bowen Qu, Olive Huang, Zeqiang Wang
arXiv AI
Jul 17

StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows

arXiv:2607. 14896v1 Announce Type: cross Abstract: Addressing a structural-engineering request requires more than a single answer; it requires a chain of interdependent artifacts: interpreted requirements, a computable model, validation records, solver outputs, code-check records, and a final report.

By Sizhong Qin, Yi Gu, Yao Jiang, Ao Cai, Changjian Zhou, Shaoxuan Shuai, Jiachang Wang, Tianhao Shen, Yueqiang Li, Xinhao Li, Li Zeng, Yueshi Chen, Dachen Gao, Genrong Xu, Wenjie Liao, Xinzheng Lu
arXiv AI
Jul 28

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications

arXiv:2607. 23124v1 Announce Type: new Abstract: Large language model agents have advanced rapidly, yet progress remains fragmented across domains, capabilities, task difficulty, and interaction settings.

By Hao Jiang, Gangtao Xin, Yingdi Huang, Guojie Zhu, Jiangshan Zhang, Xinyuan Lin, Yunkun Xu, Chengyu Shen, Wenlong Fei, Jiawei Li, Yujie Fu, Sichen Kang, Tingyu Xie, Yedi Hu, Jingren Zhang, Hongcheng Gao, Jianshu Zeng, Chong Chen, Chang Guo, Chao Feng, Feng Wang, Fulin Lin, Jinchao Ma, Lang Mei, Li Huang, Liyan Liu, Qing He, Shuting Tao, Siyu Mo, Xiangnan Chen, Xiaohan Yu, Xiaoyang Li, Yanheng Hou, Yanyu Wu, Zhihan Yang, Wentao Zhang, Yang Gao, Zhao Cao
arXiv AI
Sep 2

UniACE: A Unified Framework for Evaluating LLM Agentic Capabilities

UniACE is a unified framework that standardizes the evaluation of large language model (LLM) agents by representing each benchmark as an instruction–tool–environment triplet and running models through a shared, task‑agnostic harness in isolated runtimes. It preserves native success criteria, offers an offline mode for dynamic‑resource tasks, and standardizes efficiency metrics, execution records, and failure attribution. Applying UniACE to 7 benchmarks across 24 domains and 15 models revealed significant score shifts, ranking reversals, and sensitivity to evidence representation, highlighting the impact of evaluation configuration on reported agent performance.

By Pengyu Zhu, Lijun Li, Yaxing Lyu, Qianxin Luo, Jingyi Yang, Yi Liu, Tingfeng Hui, Xinyu Yuan, Li Sun, Sen Su, Jing Shao
arXiv AI
Jul 17

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios

arXiv:2607. 14989v1 Announce Type: cross Abstract: Large language models are increasingly evolving from text generators into general agents capable of understanding user requests, invoking external tools, and completing complex tasks through interaction.

By Chengyu Shen, Yujie Fu, Gangtao Xin, Yanheng Hou, Wenlong Fei, Guojie Zhu, Jiawei Li, Hongcheng Gao, Runming He, Zhen Hao Wong, Meiyi Qiang, Hao Liang, Zhao Cao, Hao Jiang, Chong Chen, Wentao Zhang
arXiv Machine Learning
Sep 15

Running the Gauntlet: Challenging Agentic Tasks

arXiv:2606.14397v4 Announce Type: replace Abstract: As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capab...

By Mykola Vysotskyi, Runqi Lin, Grzegorz Biziel, Michal Zakrzewski, Sebastian Montagna, Damian Rynczak, Shreyansh Padarha, Kumail Alhamoud, Zihao Fu, William Lugoloobi, Kai Rawal, Hanna Yershova, Taras Rumezhak, Guohao Li, Fazl Barez, Baoyuan Wu, Arkadiusz Drohomirecki, Yarin Gal, Chris Russell, Christopher Summerfield, Adam Mahdi, Volodymyr Karpiv, Philip Torr, Adel Bibi