arXiv AI By Benjamin Breen, Austin Letson, Borja Requena Pozo, Leopoldo Sarra

AxDafny: Agentic Verified Code Generation in Dafny

Read the original on arXiv AI →

arXiv:2606. 32007v1 Announce Type: new Abstract: We study agentic code generation in Dafny, where a model must generate both executable code and the proof artifacts for verification.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 1

SkillForge: Compositional Skill Synthesis with Verification-in-the-Loop for Generating Formally Verified Dafny Programs

SkillForge is a framework that breaks down formal code synthesis into reusable atomic skills, each handling a specific subtask such as specification inference, body synthesis, invariant generation, error diagnosis, or repair. A verification-driven harness coordinates these skills by submitting candidates to the Dafny verifier, diagnosing failures, and routing them deterministically to the appropriate repair skill until correctness is achieved or a budget is reached. On a curated benchmark, SkillForge outperforms state‑of‑the‑art agentic and iterative baselines, requiring fewer tokens and lower latency, with ablation studies showing each skill’s measurable contribution and rapid convergence.

By Yanming Liu, Xinyue Peng, Jiannan Cao, Xinyi Wang, Jinbo Su
arXiv AI
Sep 18

MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs

The paper introduces MAGS, a multi-agent framework that automatically generates executable programs with formal safety guarantees. MAGS translates LLM-generated code into the verification-aware language Dafny, repairs any safety violations using verifier feedback, and then compiles the verified code back into executable form. Evaluations on 220 diverse examples—including CUDA kernels, terminal scripts, and robotic-arm tasks—show a 100% success rate in producing programs that meet frozen safety specifications, with additional safety and functional tests confirming strong performance across domains.

By Albert Wu, Nicholas Roberts, Tzu-Heng Huang, Haoran Lin, Gil Friedman, Sungjun Cho, Gabriel Orlanski, Frederic Sala