arXiv:2608. 14615v1 Announce Type: new Abstract: Large Language Models (LLMs) perform well on established code-generation and mathematical-reasoning benchmarks, but their capabilities in mechanics and spatial geometry, here denoted as mechanical engineering awareness, has not been quantified systematically.
By Johannes Gerstmayr, Sebastian Weyrer, Tobias M\"oltner, Peter Manzl, Michael Pieber
Computer-use agents are increasingly evaluated in realistic desktop environments, but existing benchmarks provide limited coverage of professional engineering workflows whose outputs are persistent, s...
arXiv:2609.16251v1 Announce Type: new
Abstract: Computer-use agents are increasingly evaluated in realistic desktop environments, but existing benchmarks provide limited coverage of professional engi...
By Zihan Dong, Yuanzhe Liu, Zhiyuan Ma, Qishi Zhan, Dehan Kong, Guohao Li, Kaixin Li
RealCADBench is a new benchmark for evaluating intent‑to‑program parametric CAD modeling, featuring 12,632 tasks drawn from 19 factory‑automation categories and covering text, 2D drawings, product photos, and rendered images for both Part and Assembly modeling. The study reports results on a 1,770‑task evaluation slice, using metrics such as executability, Solid IoU, Surface IoU, and a rubric‑based visual‑semantic identity Judge. Across nine standalone and six frontier‑scale large models, no single model dominates all four metrics, highlighting diverse strengths and failure modes like missing fine structures and incorrect assembly placement.
By JoyIndustrial VisCAD Team, Linxin Cai, Qiuhe Hong, Zhichao Huang, Guanlin Li, Zongzhen Li, Hongsen Liu, Yichen Long, Wei Wang, Yuchen Wang, Dongyue Yang, Huimu Yu, Xianwen Zhong
arXiv:2608. 09296v1 Announce Type: new Abstract: A CAD model is not engineering-grade merely because it looks correct.
By Harmanjot Singh, Abhra Dubey, Jorge Alejandro Amador Herrera
arXiv:2606. 31252v1 Announce Type: new Abstract: Large language models can write plausible CAD scripts, but reliable industrial CAD modeling requires more than syntactically valid code: every feature, placement, and assembly relation must be accepted by an exact geometric kernel while remaining editable as parametric boundary representation geometry.
By Fumin Liu, Haoyu Zhou, Fei Hao, Lin Yang
VisCAD is a foundation model suite that tackles AI-assisted computer-aided design for industrial products, covering both part-level and assembly-level generation. Its core component, VisCAD‑M1, is a 27B model trained for part-level design generation and outperforms existing models on PubCADBench and RealCADBench, achieving a part-level score of 0.5540 and reaching 0.5797 when used as a test-time verifier. VisCAD also offers a domain-specific harness that improves complex assembly generation compared to general-purpose harnesses, showing quantitative and qualitative advantages.
By JoyIndustrial VisCAD Team, Linxin Cai, Qiuhe Hong, Zhichao Huang, Guanlin Li, Hongsen Liu, Ziqi Liu, Yichen Long, Luya Wang, Yuchen Wang, Wenxiang Wu, Huimu Yu, Ning Zhang
arXiv:2609.13234v1 Announce Type: cross
Abstract: Industrialized construction imposes stringent precision requirements on robotic assembly of modular components such as prefabricated window units. In...
By Zekai Jin, Huiguang Wang, Xiaoning Sun, Yi Shao
The paper introduces Kinematics-Grounded Agentic AI for Robotic Additive Manufacturing (A‑RAM), a framework that transforms user intent and part files into execution‑ready plans for robotic AM. It uses a large language model to interpret manufacturing goals, a deterministic Planning Agent to generate search workflows, and domain tools to evaluate slicing, placement, inverse kinematics, trajectory timing, Joint‑6 jerk, and extrusion. Experiments on a six‑axis robotic‑arm AM cell show that A‑RAM can reduce maximum Joint‑6 jerk by up to 53.5 % and mean absolute Joint‑6 jerk by 48.3 %, while also shortening motion‑plan completion times and extrusion paths.
By Jingzhan Ge, Ruimin Chen, Azadeh Haghighi, Jiong Tang, Farhad Imani
CCTU is a new benchmark designed to evaluate large language models (LLMs) on their ability to use tools under complex constraints. It includes 200 test cases that average seven constraint types and 4,700‑token prompts, covering resource, behavior, toolset, and response dimensions. An executable validation module performs step‑level checks, and nine state‑of‑the‑art LLMs were tested, revealing that none exceed a 20% task completion rate when strict constraints are enforced, with frequent violations and limited self‑refinement.
By Junjie Ye, Guoqiang Zhang, Wenjie Fu, Zelin Li, Tao Gui, Qi Zhang, Xuanjing Huang
arXiv:2607. 05573v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) and Vision-Language Models (VLMs) enable the automatic generation of parametric 3D designs from natural-language specifications.
By J de Curt\`o, Victoria Guill\'en, I. de Zarz\`a
The paper investigates the numerical reliability of gradients in differentiable physics-based optimization for robotic material manipulation. Using two Material Point Method benchmarks, it identifies three key issues: GPU many-to-one sums that alter gradient signs, reduced reliability of finite-difference checks for long rollouts, and the impact of observation and loss definitions on optimization outcomes. The study recommends reproducible accumulation, finite-difference validation, and explicit objective reporting to improve robustness in robotic optimization.