Hugging Face Trending Papers

Trustworthy Runtime Error Healing in Real-World Repositories: A Benchmark and Guardrail

arXiv AI
5d ago

Trustworthy Runtime Error Healing in Real-World Repositories: A Benchmark and Guardrail

The paper introduces HealBench, a benchmark of 265 runtime errors from 18 real-world repositories, and HealGuard, a safety framework that restricts healing code to an analyzable subset of Python and applies static and dynamic taint analysis. Using these tools, the authors evaluate a dedicated healing method and three general coding agents powered by different LLM backbones, achieving a 38.11% resume rate and a 28.68% test‑pass rate, while HealGuard flags 17.4% of successful healings as potentially unsafe. The study demonstrates that current LLM agents can meaningfully repair real repository crashes, but also highlights significant safety concerns that the Guardrail framework can detect, albeit with a high false‑positive rate.

By Gou Tan, Pengfei Chen, Zhensu Sun, Jieke Shi, Junkai Chen, Ting Zhang, Weifeng Sun, Junda He, Shuai Liang, Chuanfu Zhang, Lwin Khin Shar, David Lo
arXiv AI
Sep 4

PatchBench: Evaluating AI Agents for Vulnerability Patching

PatchBench introduces a benchmark to evaluate AI agents on realistic vulnerability patching tasks, addressing two key threats to validity: patch memorization and surface-level fixes that merely suppress crashes. The study finds that 25% of agent patches resemble historical developer patches, and that PoC-only validation inflates success rates by 1.83× on average. PatchBench mitigates these issues by selecting vulnerabilities whose true fixes lie outside the crash stack, migrating historical vulnerabilities into new contexts, and employing rigorous validation for security and semantic correctness.

By Chihao Shen, Jiacheng Li, Aastha Mahajan, Jeffery Siyuan Tian, Yonghwi Kwon, Yizheng Chen
arXiv AI
Jun 9

SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios

arXiv:2509. 22097v5 Announce Type: replace-cross Abstract: Large language model-powered code agents are rapidly transforming software engineering, yet the security risks of their generated code have become a critical concern.

By Junkai Chen, Huihui Huang, Yunbo Lyu, Junwen An, Jieke Shi, Chengran Yang, Ting Zhang, Haoye Tian, Yikun Li, Zhenhao Li, Xin Zhou, Xing Hu, David Lo
arXiv AI
Sep 24

Escaping Python Dependency Hell: A Hybrid Replay-and-Repair Pipeline for Python Dependency Resolution

The paper introduces PLLM+, a hybrid pipeline for resolving Python dependency conflicts that combines deterministic steps—such as static AST inference, replaying known successful configurations from a solutions database, and live PyPI validation—with an LLM-based repair fallback. Evaluated on the HG2.9K benchmark of 2,891 failing snippets, PLLM+ successfully fixes 1,500 cases, outperforming the baseline PLLM and reducing average runtime from 368.7 to 71.8 seconds per snippet. The majority of fixes (1,495) come from replaying existing configurations, while the LLM fallback contributes only five additional solutions.

By Veronica Poweska, Ariana Oyanguren, Jessica Pourleyli, Sourena Khanzadeh, Manar Alalfi
arXiv AI
Sep 7

When LLM Decompilers Recompile More and Preserve Less

The paper examines how large‑language‑model (LLM) decompilers, which produce clean, idiomatic C code, are currently evaluated mainly on recompilability and passing shipped tests. It shows that these metrics can mask significant behavioral differences: a decompiled function may recompile and pass all tests yet diverge on other inputs or lose disclosed vulnerabilities. To address this, the authors propose Decompile‑Diverge, a behavioral oracle that synthesizes drivers, fuzzes inputs, and compares the decompiled code’s behavior to the original, revealing divergences in up to 13% of cases and exposing gaps in current evaluation suites.

By Chang Liu, Edward Raff, Kristopher Micinski
Hugging Face Trending Papers
Aug 18

Benchmarking Automated Security Patch Backporting: How Far Are We?

The paper introduces Porting Benchmark, a dataset of 1,234 security patch backporting cases covering cross-version, cross-branch, and cross-repository scenarios, along with a unified evaluation framework. Five tools—spanning program analysis, LLM prompting, and LLM agents—are evaluated under aligned settings, revealing that performance varies significantly across tools and patch complexity, with success rates dropping sharply for structurally complex patches. The study identifies four root-cause categories for failures and demonstrates that reference-based benchmark scores may not fully reflect real-world remediation, as executable validation uncovers additional integration issues.