arXiv Computation and Language By Pierre Dantas, Lucas Cordeiro, Waldir Junior

A Benchmark Suite and Ground-Truth Methodology for Formal Verification of IEC 61131-3 Ladder Diagram Programs

Read the original on arXiv Computation and Language →

The paper introduces the first benchmark suite for formal verification of PLC programs, covering both Structured Text and Ladder Diagram encodings of IEC 61131‑3. It contains 50 programs in 83 variants across ten industrial domains, each paired with a formal property, a machine‑checkable expected verdict, and a violation witness in SV‑COMP format. The suite’s ground‑truth methodology uses construction, fault injection, and cross‑tool consensus to ensure reliable verdicts, and the authors validate it with ESBMC and nuXmv, revealing format and semantics fragmentation in existing tools.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
4d ago

From Dead Code and Static Requirements to Working Engines: Software Revival with Coding Agents

The paper introduces ReviveBench, a benchmark designed to evaluate coding agents’ ability to revive non‑running software and reconstruct industrial engines from open specifications. It comprises two families of tasks—revival (ten tasks addressing dependency issues, missing modules, legacy builds, and GPU models) and reconstruction (thirteen tasks covering numerical, geometric, hardware, and transactional systems). The benchmark uses hidden verifiers calibrated against native environments, engineering tools, or reference implementations, and the authors report that the strongest evaluated model passes all revival tasks and most reconstruction tasks, while also uncovering verifier defects that highlight measurement error in executable verification.

By Tianyu Liu, Dingyuan Dai, Yufan Du, Zhen Yang