arXiv AI By Ali ArjomandBigdeli, Jiawei Zhou, Stanley Bak

Large Language Models as Falsifiers for Cyber-Physical Systems

Read the original on arXiv AI →

The paper introduces LLM-Falsifier, a large language model–based method for falsifying cyber‑physical system specifications written in Signal Temporal Logic (STL). By exposing the LLM to semantic cues such as natural‑language names, output trajectories, and critical‑time witnesses, the approach performs smarter, sample‑efficient robustness searches. On ARCH‑COMP benchmarks, LLM‑Falsifier outperforms existing tools across 14 of 21 specifications, requiring fewer simulations to find counterexamples.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Sep 17

Large Language Models as Falsifiers for Cyber-Physical Systems

The paper introduces LLM-Falsifier, a method that uses large language models to find counterexamples in cyber‑physical systems by minimizing Signal Temporal Logic robustness. By providing the LLM with natural‑language context, trajectory outputs, and critical‑time witnesses, the approach achieves smarter, more sample‑efficient searches. On ARCH‑COMP benchmarks, LLM‑Falsifier outperforms existing tools across 14 of 21 specifications in terms of simulations needed to locate a counterexample.

arXiv Computation and Language
Sep 10

$\Phi$-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?

arXiv:2609.10226v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable capabilities in reasoning and code generation, raising the prospect that they could assist in...

By Leilei Ding, Shumin Wang, Yuting Huang, Fanqi Wan, Yinmin Zhang, Qi Han, Yiming Xu, Feiyuan Zhang, Xiaomeng Chu, Guoliang You, Wuyang Zhang, Daxin Jiang, Yanyong Zhang
arXiv AI
Aug 18

ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization

arXiv:2602. 15983v3 Announce Type: replace-cross Abstract: Large language models (LLMs) can translate natural language into optimization code, but silent failures pose a critical risk: code that executes and returns solver-feasible solutions may encode semantically incorrect formulations---a feasibility--correctness gap reaching 90 percentage points on compositional problems.

By Junbo Jacob Lian, Yujun Sun, Huiling Chen, Chaoyu Zhang, Hanzhang Qin, Chung-Piaw Teo
arXiv AI
Aug 5

IR2Solve: Structured Intermediate Representations for Cost-Efficient Optimization Autoformulation

arXiv:2608. 02641v1 Announce Type: cross Abstract: Large language models (LLMs) can translate natural-language optimization problems into solver-ready formulations, but direct code generation is brittle: schema, indexing, and semantic errors can cause compilation failures, infeasible models, or incorrect objectives, while iterative repair, search, and multi-agent workflows increase inference cost.

By Penglin Zhu, Linhai Zhang, Jungang Xu, Xinchi Wei, Xiuqi Wu