arXiv:2608.28725v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly used as graders, verifiers, and process auditors, but most mathematical evaluations still emphasize final...
By Fateme Mazdarani, Carlos Toxtli
arXiv:2602. 15983v3 Announce Type: replace-cross Abstract: Large language models (LLMs) can translate natural language into optimization code, but silent failures pose a critical risk: code that executes and returns solver-feasible solutions may encode semantically incorrect formulations---a feasibility--correctness gap reaching 90 percentage points on compositional problems.
By Junbo Jacob Lian, Yujun Sun, Huiling Chen, Chaoyu Zhang, Hanzhang Qin, Chung-Piaw Teo
arXiv:2606. 06133v1 Announce Type: cross Abstract: TLA+ is a formal specification language for verifying distributed systems and safety-critical protocols.
By Eric Spencer, Arslan Bisharat, Brian Ortiz, Khushboo Bhadauria, TaiNing Wang, George K. Thiruvathukal, Konstantin Laufer, Mohammed Abuhamad
The paper introduces LLM-Falsifier, a large language model–based method for falsifying cyber‑physical system specifications written in Signal Temporal Logic (STL). By exposing the LLM to semantic cues such as natural‑language names, output trajectories, and critical‑time witnesses, the approach performs smarter, sample‑efficient robustness searches. On ARCH‑COMP benchmarks, LLM‑Falsifier outperforms existing tools across 14 of 21 specifications, requiring fewer simulations to find counterexamples.
By Ali ArjomandBigdeli, Jiawei Zhou, Stanley Bak
arXiv:2609.21190v1 Announce Type: cross
Abstract: Ensuring the correctness of LLM-generated code is a core challenge for modern software engineering. Benchmarks for agentic code generation check corr...
By George Ma, Benjamin Mikek, Haoyu Li, Ferhat Erata, Yuhao Zhang, Zeren Shui, Behrooz Omidvar Tehrani, Jun Huan, Murali Krishna Ramanathan, Somayeh Sojoudi, Hao Zhou, Anoop Deoras
arXiv:2609.35841v1 Announce Type: cross
Abstract: Mutation testing evaluates test-suite adequacy by injecting synthetic faults into program code. However, traditional rule-based tools often generate...
By Nils Kiele, Zainab Saad, Zirui Wang, Steve Drew, Samira Ebrahimi Kahou