arXiv AI

Djinnlang: Higher-Level Programming by Unambiguous Specification with an LLM in the Compiler

arXiv AI
Sep 11

Spec-Harness: Measuring and Improving Behavioral Adequacy of LLM-Synthesized Formal Specifications

Spec‑Harness evaluates how well large language models (LLMs) synthesize Java Modeling Language (JML) specifications by measuring behavioral adequacy across precondition and postcondition correctness and completeness. The study shows that while prompt optimization can raise verifier pass rates, many accepted specifications remain behaviorally weak, either over‑ or under‑constraining inputs and outputs. Spec‑Harness also serves as a feedback mechanism that improves the quality of specifications generated by general‑purpose coding agents and a specialized JML agent.

By Md Rakib Hossain Misu, Iris Ma, Cristina V. Lopes
Hugging Face Trending Papers
Jul 7

Harnessing Code Agents for Automatic Software Verification

Formal verification offers the strongest guarantee of software correctness, but it does not scale: the proofs demanded by interactive theorem provers such as Coq require enormous expert effort. Large language models (LLMs) promise to generate these proofs automatically, yet existing approaches wire a fixed, human-designed proof strategy into the system and constrain the model to follow it (retrieving premises and predicting tactics one step at a time, or splitting goals by divide-and-conquer), and still prove only a fraction of their target theorems.

arXiv AI
4d ago

Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity

The paper introduces JAZ, a minimalist LLM agent framework that centers on a single primitive called “invoke”, which allows an LLM to write and execute arbitrary code, including recursive calls, while treating all inputs and interaction history as variables in the code environment. JAZ provides built‑in hooks for constraints and monitoring but relies solely on prompting, without external tools, memory systems, or file‑system access. Experiments show that JAZ “invoke” outperforms specialized external harnesses such as Letta (MemGPT) and ACE on long‑horizon recall tasks and continual self‑improvement, achieving higher accuracy at lower cost.

By Zhening Li, Joshua Liu, Mateja Vukelic, Nicole Shen, Supriya Lall, Amitayush Thakur, Alex Zhang, Omar Khattab, Jonathan Light, Armando Solar-Lezama