arXiv AI

The Illusion of Secure LLM Code: Closing the Security Gap via Iterative Reprompting

arXiv:2607. 23710v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly integrated into software development workflows, yet their ability to autonomously generate secure authentication code remains uncertain.

arXiv AI
Aug 5

AgenticSCR: An Autonomous Agentic Secure Code Review for Immature Vulnerabilities Detection

arXiv:2601. 19138v2 Announce Type: replace-cross Abstract: Secure code review is critical during pre-integration, where Atlassian developers rely on lightweight analysis tools, while deep security assessment is deferred to later stages, delaying feedback and increasing remediation costs.

By Wachiraphan Charoenwet, Kla Tantithamthavorn, Patanamon Thongtanunam, Hong Yi Lin, Minwoo Jeong, Ming Wu
arXiv Computation and Language
Sep 15

Toward Secure Code Generation: Bridging Correctness and Security via Task-Adaptive Vulnerability Modeling and Execution-Based Benchmarking

arXiv:2407.02395v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used for program synthesis, yet they often generate code that is functionally plausible but ins...

By Jiexin Wang, Liuwen Cao, Xitong Luo, Yang Cao, Zhenghao Li, Yunyi Xiao, Mengchen Zhao, Adam Jatowt, Yi Cai
arXiv Computation and Language
3d ago

SecureVibe: Making Vibe Coding More Secure

SecureVibe is a training recipe designed to enhance the security of vibe coding by explicitly targeting planning and testing for code security. It combines supervised fine‑tuning on a security suite with post‑training methods (SECUREVIBE_rl and SECUREVIBE_hg) that use verifiable execution feedback and hint‑based self‑supervision. The approach outperforms baselines on multiple security coding benchmarks, improving security pass@1 by up to 6.9 points on BaxBench and 11.5 points on SusVibes, while also boosting functionality pass@1 on both security and generic coding tasks.

By Danqing Wang, Baolin Peng, Zhepei Wei, Isadora White, Wenlin Yao, Hao Cheng, Qianhui Wu, Minseon Kim, Xingdi Yuan, Lei Li, Jianfeng Gao
arXiv AI
Aug 24

Vibe Coding and Web Application Security: A Twin-Prompt Study

The study examines how adding a security-requirements section to prompts affects web applications generated by a large language model. Six distinct applications were produced twice—once with a baseline prompt and once with a security-aware prompt—yielding 12 programs. Analysis of these programs revealed 75 confirmed security findings, with the security-aware variants showing fewer issues (24 vs. 51) and no Critical or High severity problems.

By Darko Andro\v{c}ec
arXiv AI
Jun 16

DualGauge: Automated Joint Security-Functionality Benchmarking of Specification-Only Code Generation by LLMs and Coding Agents

arXiv:2511. 20709v2 Announce Type: replace-cross Abstract: Large language models (LLMs) and LLM-based coding agents are now used to generate code from natural-language specifications, yet ensuring such code is both functionally correct and secure remains a challenge.

By Rupam Patir, Keyan Guo, Suvadra Barua, Abhijeet Pathak, Dinesh Gudimetla, Jiawei Guo, Hongxin Hu, Haipeng Cai