arXiv:2606. 10928v1 Announce Type: cross Abstract: Large language models can reduce the manual effort required to set up finite element simulations, but they introduce reliability risks when generated solver code lies on the critical path.
By Nilay Upadhyay, Wesley F. Reinhart
arXiv:2609.13152v1 Announce Type: new
Abstract: Large language models (LLMs) perform strongly on static science benchmarks, yet their ability to reason about the physical world through active experim...
By Joseph Chan, Utkarsh Jha, Xiyin Yang, Abhinav Jarajapu, Anik Sahai, Eddie Hu, Robin Jeshua Deepak, Stefano Saravalle, Aditya Shah
arXiv:2605. 28579v2 Announce Type: replace Abstract: Large language models (LLMs) have recently advanced text-driven 3D generation, yet Text-to-CAD remains far from supporting industrial product design.
By Xiaoyu Dong, Zhi Li, Xiao-Ming Wu
arXiv:2607. 05573v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) and Vision-Language Models (VLMs) enable the automatic generation of parametric 3D designs from natural-language specifications.
By J de Curt\`o, Victoria Guill\'en, I. de Zarz\`a
arXiv:2606. 31252v1 Announce Type: new Abstract: Large language models can write plausible CAD scripts, but reliable industrial CAD modeling requires more than syntactically valid code: every feature, placement, and assembly relation must be accepted by an exact geometric kernel while remaining editable as parametric boundary representation geometry.
By Fumin Liu, Haoyu Zhou, Fei Hao, Lin Yang
SCICONVBENCH is a benchmark designed to evaluate large language models (LLMs) on multi‑turn clarification tasks in computational science. It focuses on two key abilities: eliciting missing information (disambiguation) and resolving contradictory requests (inconsistency resolution) across four domains—fluid mechanics, solid mechanics, materials science, and partial differential equations. The benchmark pairs a structured task ontology with a rubric‑based evaluation framework, measuring LLM performance in clarification behavior, conversational grounding, and final‑specification fidelity, and reveals that even top models only resolve about 52.7% of disambiguation cases in fluid mechanics while often making ungrounded assumptions.
By Nithin Somasekharan, Youssef Hassan, Shiyao Lin, Gihan Panapitiya, Patrick Emami, Anurag Acharya, Sameera Horawalavithana, Shaowu Pan
arXiv:2609.22609v1 Announce Type: cross
Abstract: Insertion is a fundamental operation in robotic construction assembly, where variations in material properties and assembly conditions make it diffic...
By Lin He, Yanyi Chen, Haofei Sun, Lingyao Li, Min Deng
arXiv:2609.21493v1 Announce Type: new
Abstract: Multimodal large language models, or MLLMs, perform well at visual understanding and structured generation, yet these capabilities do not establish whe...
By Zicheng Zhao, Dongyin Chen, Rui Xu, Yinghui Xu
arXiv:2608. 09296v1 Announce Type: new Abstract: A CAD model is not engineering-grade merely because it looks correct.
By Harmanjot Singh, Abhra Dubey, Jorge Alejandro Amador Herrera
arXiv:2503.18460v2 Announce Type: replace-cross
Abstract: Modelica is a widely adopted language for simulating complex physical systems, yet effective model creation and optimization require substant...
By Jiahui Xiang, Tong Ye, Peiyu Liu, Yinan Zhang, Wenhai Wang
arXiv:2606. 00138v1 Announce Type: new Abstract: Finite element analysis (FEA) is the most important numerical approach for solid mechanics.
By Titu Ranjan Sarker, Muhammed Jawaad Zulqernine, Ling Yue, Shaowu Pan, Chenxi Wang, Shiyao Lin
MechReason is a new benchmark for multi-image, multi-hop reasoning in mechanical engineering, featuring 12,000 question-answer pairs with explicit reasoning chains and 21,000 visual items across nine evidence types. It covers eight task types—explanation, prediction, design, and diagnosis—within four reasoning dimensions, and is constructed through a four-stage pipeline that extracts engineering claims, generates verification-masked questions, and validates multimodal quality. Even state‑of‑the‑art models score only 62.89% on this dataset, highlighting its difficulty.
By Tengyue Wang, Kang An, Chenxu Du, Zhongyu Yang, Yuanchi Zhu, Xinqi Yang, Hebao Zhu, Ziliang Wang, FaQiang Qian, Yunli Yang, Qibing Ren