Trustworthy Runtime Error Healing in Real-World Repositories: A Benchmark and Guardrail
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The paper introduces HealBench, a benchmark of 265 runtime errors from 18 real-world repositories, and HealGuard, a safety framework that restricts healing code to an analyzable subset of Python and applies static and dynamic taint analysis. Using these tools, the authors evaluate a dedicated healing method and three general coding agents powered by different LLM backbones, achieving a 38.11% resume rate and a 28.68% test‑pass rate, while HealGuard flags 17.4% of successful healings as potentially unsafe. The study demonstrates that current LLM agents can meaningfully repair real repository crashes, but also highlights significant safety concerns that the Guardrail framework can detect, albeit with a high false‑positive rate.
arXiv:2605. 17450v2 Announce Type: replace-cross Abstract: As software systems grow increasingly complex, automated vulnerability repair (AVR) remains difficult because the materials available to a repair system are usually failure artifacts rather than repair guidance.
PatchBench introduces a benchmark to evaluate AI agents on realistic vulnerability patching tasks, addressing two key threats to validity: patch memorization and surface-level fixes that merely suppress crashes. The study finds that 25% of agent patches resemble historical developer patches, and that PoC-only validation inflates success rates by 1.83× on average. PatchBench mitigates these issues by selecting vulnerabilities whose true fixes lie outside the crash stack, migrating historical vulnerabilities into new contexts, and employing rigorous validation for security and semantic correctness.
arXiv:2504. 20412v3 Announce Type: replace-cross Abstract: Fuzzing frameworks like syzkaller have uncovered thousands of Linux kernel crashes, many of which are critical and security-sensitive.
arXiv:2607. 11098v1 Announce Type: cross Abstract: Tool-using LLM agents are mostly evaluated assuming all tools work.
arXiv:2606. 30755v1 Announce Type: cross Abstract: Claw-like AI agents (e.