arXiv AI By Haifeng Li, Mo Hai

Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits

Read the original on arXiv AI →

arXiv:2607. 16646v1 Announce Type: cross Abstract: Large language models now translate natural-language descriptions of decision problems into solver-ready optimization models, but they fail silently.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 18

ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization

arXiv:2602. 15983v3 Announce Type: replace-cross Abstract: Large language models (LLMs) can translate natural language into optimization code, but silent failures pose a critical risk: code that executes and returns solver-feasible solutions may encode semantically incorrect formulations---a feasibility--correctness gap reaching 90 percentage points on compositional problems.

By Junbo Jacob Lian, Yujun Sun, Huiling Chen, Chaoyu Zhang, Hanzhang Qin, Chung-Piaw Teo
arXiv AI
Sep 18

Large Language Models as Falsifiers for Cyber-Physical Systems

The paper introduces LLM-Falsifier, a large language model–based method for falsifying cyber‑physical system specifications written in Signal Temporal Logic (STL). By exposing the LLM to semantic cues such as natural‑language names, output trajectories, and critical‑time witnesses, the approach performs smarter, sample‑efficient robustness searches. On ARCH‑COMP benchmarks, LLM‑Falsifier outperforms existing tools across 14 of 21 specifications, requiring fewer simulations to find counterexamples.

By Ali ArjomandBigdeli, Jiawei Zhou, Stanley Bak
arXiv AI
Sep 21

SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?

arXiv:2609.21190v1 Announce Type: cross Abstract: Ensuring the correctness of LLM-generated code is a core challenge for modern software engineering. Benchmarks for agentic code generation check corr...

By George Ma, Benjamin Mikek, Haoyu Li, Ferhat Erata, Yuhao Zhang, Zeren Shui, Behrooz Omidvar Tehrani, Jun Huan, Murali Krishna Ramanathan, Somayeh Sojoudi, Hao Zhou, Anoop Deoras