Code-Augur: Agentic Vulnerability Detection via Specification Inference
arXiv:2606. 18619v1 Announce Type: cross Abstract: The advent of agentic vulnerability detection is already becoming a watershed moment for software security.
arXiv:2603. 26270v2 Announce Type: replace-cross Abstract: Smart contracts govern billions of dollars in decentralized finance (DeFi), yet automated vulnerability detection remains challenging because many vulnerabilities are tightly coupled with project-specific business logic.
arXiv:2606. 18619v1 Announce Type: cross Abstract: The advent of agentic vulnerability detection is already becoming a watershed moment for software security.
arXiv:2606. 03128v1 Announce Type: cross Abstract: Smart contracts face critical security challenges that require thorough auditing in decentralized web services.
arXiv:2606. 26216v1 Announce Type: cross Abstract: We present CyberChainBench, a benchmark for evaluating LLM-based agents on smart contract security across three complementary tasks: vulnerability detection, exploit generation, and patch synthesis.
AgentXploit is a two‑role auditing system that separates repository‑level attack‑path discovery from runtime exploitation for AI agents. The Analyzer Agent traces attacker‑controlled inputs to sensitive operations and records candidate attack paths, while the Exploiter Agent turns these paths into concrete attacks and refines them using runtime feedback. The system is evaluated on AgentXploit‑Bench, a benchmark of 72 reproducible vulnerabilities across 12 open‑source AI‑agent systems, achieving 59.3% end‑to‑end success compared to 38.4% for Codex, and 79.2% attack success on AgentDojo versus 52.7% for AgentVigil.
arXiv:2607. 11698v1 Announce Type: cross Abstract: Production LLM agents such as Claude Code and Codex operate over untrusted content, files, commands, and workspace state, making safety failures directly actionable.
Production LLM agents such as Claude Code and Codex operate over untrusted content, files, commands, and workspace state, making safety failures directly actionable. Red-teaming must therefore keep pace with evolving models and tools.
arXiv:2601. 19138v2 Announce Type: replace-cross Abstract: Secure code review is critical during pre-integration, where Atlassian developers rely on lightweight analysis tools, while deep security assessment is deferred to later stages, delaying feedback and increasing remediation costs.
arXiv:2606. 13757v1 Announce Type: cross Abstract: Large language model (LLM) reviewers are increasingly used in pull-request (PR) workflows, where their approvals help decide which code is merged into a repository.
arXiv:2609.08040v1 Announce Type: cross Abstract: The software supply chain has become an increasingly exposed attack surface because of its reliance on intricate yet fragile dependencies. Existing d...
arXiv:2607. 08288v1 Announce Type: cross Abstract: In critical infrastructure, operational technology environments often cannot be actively scanned, and yet active system feedback is needed for risk assessment and compliance.
PatchBench introduces a benchmark to evaluate AI agents on realistic vulnerability patching tasks, addressing two key threats to validity: patch memorization and surface-level fixes that merely suppress crashes. The study finds that 25% of agent patches resemble historical developer patches, and that PoC-only validation inflates success rates by 1.83× on average. PatchBench mitigates these issues by selecting vulnerabilities whose true fixes lie outside the crash stack, migrating historical vulnerabilities into new contexts, and employing rigorous validation for security and semantic correctness.
The paper presents a method that employs Large Language Models to automatically inject known vulnerabilities into Solidity smart contracts. Using a multi-step validation pipeline, the authors generate nearly 1,000 candidate contracts from real-world sources, ultimately confirming 32 vulnerable variants across 25 vulnerability types. These validated contracts are then used to evaluate the coverage of three static analysis tools, highlighting both complementary strengths and gaps in current detection approaches.