arXiv AI

CREBench: Evaluating Large Language Models in Cryptographic Binary Reverse Engineering

arXiv:2604. 03750v2 Announce Type: replace-cross Abstract: Reverse engineering (RE) is central to software security, particularly for cryptographic programs that handle sensitive data and are highly prone to vulnerabilities.

arXiv AI
Aug 13

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark

arXiv:2608. 11469v1 Announce Type: cross Abstract: AI agents are rapidly improving in cybersecurity capabilities when the source code is available for analysis, yet much of the software most consequential to cybersecurity, including malware, firmware, and proprietary applications, is available only as binaries.

By Jeremy Spence, Nicholas Assaderaghi, Jinhao Zhu, Nikil Ravi, Raluca Ada Popa, Guannan Wei, Yangruibo Ding, Zhuo Zhang
arXiv AI
Jun 11

MPC-Patch-Bench: Security-Aware LLM Code Patch for Multi-Party Computation

arXiv:2606. 11416v1 Announce Type: cross Abstract: Repository-level benchmarks for evaluating Large Language Model (LLM) code repair on Secure Multi-Party Computation (MPC) software do not yet exist, and directly transplanting general-purpose benchmarks such as SWE-bench fails on three structural fronts: (i) MPC repositories are dominated by generic Python infrastructure rather than cryptographic logic; (ii) high-value MPC fixes lack the standardized tests rigid extraction pipelines require; and (iii) standard fail-to-pass evaluation is insufficient for code that must also be cryptographically safe.

By Yukuan Zhang, Mengxin Zheng, Qian Lou
arXiv AI
4d ago

Cheap to Hypothesize, Costly to Verify: The Defense Surface of Agentic Vulnerability Discovery

The paper introduces RedHerring, a defense mechanism that inserts safe decoy vulnerabilities into code repositories to divert autonomous LLM agents’ verification efforts away from real security flaws. By embedding CVE-derived vulnerability chains with false bridges and providing a private certificate for quick verification, RedHerring forces agents to spend a significant portion of their limited resources on decoys. Experiments on 33 OSS‑Fuzz projects show a 38.7‑60.4% reduction in discovered real vulnerabilities, even when agents are aware of decoys.

By Kaikai Zhang, Zihan Zhang, Yuchong Xie, Zesen Liu, Shuangjie Yao, Zhixiang Zhang, Dongdong She
arXiv AI
Sep 21

CESBench: Benchmarking Large Language Models on Cryptographic Engineering Security for IoT Devices

CESBench is a new benchmark for evaluating large language models on cryptographic engineering security for IoT devices, comprising 380 expert‑written items across six sub‑domains such as side‑channel, fault injection, and implementation. The benchmark includes four task types—multiple‑choice, judgment, scenario, and code—each designed to test different competences, with automatic scoring for the first two and LLM‑based judging for the latter two. Evaluation of 11 open‑weight and proprietary LLMs shows strong performance on multiple‑choice and code tasks but weaker results on judgment and scenario tasks, highlighting gaps in justifying security verdicts.

By Wenquan Zhou, An Wang, Jing Liang, Peien Feng, Jingqi Zhang, Yaoling Ding, Liehuang Zhu
arXiv Computation and Language
Sep 15

Toward Secure Code Generation: Bridging Correctness and Security via Task-Adaptive Vulnerability Modeling and Execution-Based Benchmarking

arXiv:2407.02395v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used for program synthesis, yet they often generate code that is functionally plausible but ins...

By Jiexin Wang, Liuwen Cao, Xitong Luo, Yang Cao, Zhenghao Li, Yunyi Xiao, Mengchen Zhao, Adam Jatowt, Yi Cai
arXiv AI
Jun 9

SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios

arXiv:2509. 22097v5 Announce Type: replace-cross Abstract: Large language model-powered code agents are rapidly transforming software engineering, yet the security risks of their generated code have become a critical concern.

By Junkai Chen, Huihui Huang, Yunbo Lyu, Junwen An, Jieke Shi, Chengran Yang, Ting Zhang, Haoye Tian, Yikun Li, Zhenhao Li, Xin Zhou, Xing Hu, David Lo
arXiv Machine Learning
Jul 8

Multi-Channel Spread-Spectrum Code Watermarking

arXiv:2607. 06009v1 Announce Type: cross Abstract: Attributing code to the large language model that produced it is essential for provenance, licensing, and misuse accountability, yet no deployed watermark meets this need.

By Soohyeon Choi, Debin Gao, Yue Duan
arXiv AI
Sep 11

CS-Guard: Benchmarking LLM Guardrails for Code Generation Security

CS-Guard is a new benchmark that systematically evaluates guardrails for code generation security, covering 1,000 malware-generation prompts, 7 jailbreak attacks, and a novel fictional scenario attack (FSA) for text-to-code generation, as well as 331 code prompts for code-to-code generation. The study empirically tests nine guardrails across seven large language models, finding that many guardrails fail to prevent malicious code generation, with attack success rates reaching about 50% for text-to-code and up to nearly 100% for code-to-code and FSA scenarios. CS-Guard introduces a modular three-layer guardrail taxonomy and releases its benchmark and data to support future research.

By Jinyang Li, Mingyu Guo, Hung X. Nguyen