arXiv AI

Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits

arXiv:2607. 16646v1 Announce Type: cross Abstract: Large language models now translate natural-language descriptions of decision problems into solver-ready optimization models, but they fail silently.

arXiv AI
Aug 18

ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization

arXiv:2602. 15983v3 Announce Type: replace-cross Abstract: Large language models (LLMs) can translate natural language into optimization code, but silent failures pose a critical risk: code that executes and returns solver-feasible solutions may encode semantically incorrect formulations---a feasibility--correctness gap reaching 90 percentage points on compositional problems.

By Junbo Jacob Lian, Yujun Sun, Huiling Chen, Chaoyu Zhang, Hanzhang Qin, Chung-Piaw Teo
arXiv AI
Sep 18

Large Language Models as Falsifiers for Cyber-Physical Systems

The paper introduces LLM-Falsifier, a large language model–based method for falsifying cyber‑physical system specifications written in Signal Temporal Logic (STL). By exposing the LLM to semantic cues such as natural‑language names, output trajectories, and critical‑time witnesses, the approach performs smarter, sample‑efficient robustness searches. On ARCH‑COMP benchmarks, LLM‑Falsifier outperforms existing tools across 14 of 21 specifications, requiring fewer simulations to find counterexamples.

By Ali ArjomandBigdeli, Jiawei Zhou, Stanley Bak
arXiv AI
Sep 21

SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?

arXiv:2609.21190v1 Announce Type: cross Abstract: Ensuring the correctness of LLM-generated code is a core challenge for modern software engineering. Benchmarks for agentic code generation check corr...

By George Ma, Benjamin Mikek, Haoyu Li, Ferhat Erata, Yuhao Zhang, Zeren Shui, Behrooz Omidvar Tehrani, Jun Huan, Murali Krishna Ramanathan, Somayeh Sojoudi, Hao Zhou, Anoop Deoras
Hugging Face Trending Papers
Sep 17

Large Language Models as Falsifiers for Cyber-Physical Systems

The paper introduces LLM-Falsifier, a method that uses large language models to find counterexamples in cyber‑physical systems by minimizing Signal Temporal Logic robustness. By providing the LLM with natural‑language context, trajectory outputs, and critical‑time witnesses, the approach achieves smarter, more sample‑efficient searches. On ARCH‑COMP benchmarks, LLM‑Falsifier outperforms existing tools across 14 of 21 specifications in terms of simulations needed to locate a counterexample.

arXiv AI
Aug 3

ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models

arXiv:2607. 29431v1 Announce Type: new Abstract: Large language models increasingly generate optimization models from natural language, but existing evaluation often reduces a generated model and its ground truth to a single equivalent/not-equivalent verdict or an execution-success rate--labels that are neither independently checkable nor faithful to the multiple distinct senses in which two formulations can agree.

By Penglin Zhu, Jungang Xu