Learning from Failures: A Failure-Driven Prompt Refinement for LLM-Based Vulnerability Analysis
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper introduces VERA, an automated framework that audits large language model (LLM) reasoning in software vulnerability analysis. Instead of relying on free‑form explanations, VERA requires models to produce a Structured Reasoning Record (SRR) that captures pointers, memory operations, and state transitions in machine‑readable fields. A multi‑stage judge then checks each SRR against eight reasoning failure modes, revealing that reasoning flaws are as common in correct verdicts as in incorrect ones and that VERA detects 87% of errors missed by free‑form LLM‑as‑judge evaluations.
Software vulnerability remediation is a cognitively demanding task that requires specialized security expertise often lacking in general developers. In the meantime, Large Language Models (LLMs) assisted tools show potential in vulnerability detection, location, and repair tasks.
arXiv:2606. 03601v1 Announce Type: cross Abstract: While safety alignment and guardrails help large language models (LLMs) avoid harmful outputs, they can also induce overrefusal, i.
arXiv:2607. 05842v1 Announce Type: cross Abstract: Large language model (LLM)-assisted software security operates at a difficult boundary: the vulnerability-analysis terminology needed for legitimate code review, triage, and repair can closely resemble terminology associated with misuse.
arXiv:2606. 30587v1 Announce Type: cross Abstract: Researchers and practitioners increasingly apply Large Language Models (LLMs) for automated vulnerability detection.
arXiv:2607. 03833v1 Announce Type: cross Abstract: While Large Language Models (LLMs) have achieved remarkable success in Text-to-SQL tasks, their deployment in real-world environments is hindered by latent reliability issues.