arXiv Computation and Language

A Benchmark Suite and Ground-Truth Methodology for Formal Verification of IEC 61131-3 Ladder Diagram Programs

The paper introduces the first benchmark suite for formal verification of PLC programs, covering both Structured Text and Ladder Diagram encodings of IEC 61131‑3. It contains 50 programs in 83 variants across ten industrial domains, each paired with a formal property, a machine‑checkable expected verdict, and a violation witness in SV‑COMP format. The suite’s ground‑truth methodology uses construction, fault injection, and cross‑tool consensus to ensure reliable verdicts, and the authors validate it with ESBMC and nuXmv, revealing format and semantics fragmentation in existing tools.

arXiv AI
4d ago

From Dead Code and Static Requirements to Working Engines: Software Revival with Coding Agents

The paper introduces ReviveBench, a benchmark designed to evaluate coding agents’ ability to revive non‑running software and reconstruct industrial engines from open specifications. It comprises two families of tasks—revival (ten tasks addressing dependency issues, missing modules, legacy builds, and GPU models) and reconstruction (thirteen tasks covering numerical, geometric, hardware, and transactional systems). The benchmark uses hidden verifiers calibrated against native environments, engineering tools, or reference implementations, and the authors report that the strongest evaluated model passes all revival tasks and most reconstruction tasks, while also uncovering verifier defects that highlight measurement error in executable verification.

By Tianyu Liu, Dingyuan Dai, Yufan Du, Zhen Yang
arXiv AI
Sep 21

SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?

arXiv:2609.21190v1 Announce Type: cross Abstract: Ensuring the correctness of LLM-generated code is a core challenge for modern software engineering. Benchmarks for agentic code generation check corr...

By George Ma, Benjamin Mikek, Haoyu Li, Ferhat Erata, Yuhao Zhang, Zeren Shui, Behrooz Omidvar Tehrani, Jun Huan, Murali Krishna Ramanathan, Somayeh Sojoudi, Hao Zhou, Anoop Deoras
arXiv AI
Sep 15

Natural-Language to SysMLv2 Translation via Conformance-Driven Iterative Refinement

The paper introduces a framework that translates natural‑language descriptions into SysMLv2 models using a generate‑check‑repair loop driven by a SysMLv2 conformance checker. By embedding the checker as an oracle, the system iteratively repairs generated models until they achieve zero conformance errors, ensuring they are deployable in industrial modeling environments. Evaluation on 151 prompts across four large language models shows the approach raises production‑conformance acceptance from 51.16% to 100%.

By Chance LaVoie, Eladio Andujar Lugo, Taylan G. Topcu, Levent Burak Kara
arXiv AI
Jul 24

Towards a Certifying Grounder

arXiv:2607. 21199v1 Announce Type: cross Abstract: Grounding, the translation of high-level theories into equivalent quantifier-free formulas, is a crucial step in declarative solving, yet it has so far escaped the proof-logging revolution.

By Daimy Van Caudenberg, Alexander Ek, Carlos Cantero, Bart Bogaerts
arXiv AI
3d ago

Self-Spec Verifiable Code Generation

arXiv:2609.39568v1 Announce Type: cross Abstract: Large language models (LLMs) may generate unreliable code on corner cases missed by testing, while formal verification can provide machine-checkable...

By Jiaru Qian, Yihong Dong, Yongmin Li, Hao Zhu, Bin Gu, Ge Li