arXiv AI By Yunji Kim, Yunseok Lee, Hyunwoo Seo, Jaerim Choi, Woojin Lee

PLCWorld: Benchmarking LLM-Generated PLC Programs in Closed-Loop Plant Simulation

Read the original on arXiv AI →

PLCWorld is a closed‑loop execution environment and benchmark that evaluates large language model (LLM)–generated programmable logic controller (PLC) programs by coupling Structured Text (ST) execution with simulated plant responses and sensor feedback. It includes 100 synthetic tasks and 473 task‑condition pairs across Motion Control and Material Handling, with difficulty levels based on control‑dependency scope. The benchmark reports Task Success and Safety Violation separately and validates results through practitioner review, reference programs, counterexamples, and comparisons with independent ST runtimes, revealing performance differences among six LLMs and four generation‑and‑verification workflows.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 17

A Benchmark Suite and Ground-Truth Methodology for Formal Verification of IEC 61131-3 Ladder Diagram Programs

The paper introduces the first benchmark suite for formal verification of PLC programs, covering both Structured Text and Ladder Diagram encodings of IEC 61131‑3. It contains 50 programs in 83 variants across ten industrial domains, each paired with a formal property, a machine‑checkable expected verdict, and a violation witness in SV‑COMP format. The suite’s ground‑truth methodology uses construction, fault injection, and cross‑tool consensus to ensure reliable verdicts, and the authors validate it with ESBMC and nuXmv, revealing format and semantics fragmentation in existing tools.

By Pierre Dantas, Lucas Cordeiro, Waldir Junior
arXiv Computation and Language
Sep 22

CCTU: A Benchmark for Tool Use under Complex Constraints

CCTU is a new benchmark designed to evaluate large language models (LLMs) on their ability to use tools under complex constraints. It includes 200 test cases that average seven constraint types and 4,700‑token prompts, covering resource, behavior, toolset, and response dimensions. An executable validation module performs step‑level checks, and nine state‑of‑the‑art LLMs were tested, revealing that none exceed a 20% task completion rate when strict constraints are enforced, with frequent violations and limited self‑refinement.

By Junjie Ye, Guoqiang Zhang, Wenjie Fu, Zelin Li, Tao Gui, Qi Zhang, Xuanjing Huang
arXiv AI
Jul 1

Embodied CAD: Solver-Grounded LLM Agents for Parametric B-Rep Assembly Modeling

arXiv:2606. 31252v1 Announce Type: new Abstract: Large language models can write plausible CAD scripts, but reliable industrial CAD modeling requires more than syntactically valid code: every feature, placement, and assembly relation must be accepted by an exact geometric kernel while remaining editable as parametric boundary representation geometry.

By Fumin Liu, Haoyu Zhou, Fei Hao, Lin Yang