arXiv Computation and Language
Aug 24

Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks

The paper introduces SUSVIBES, a benchmark of 186 real‑world software engineering tasks where human programmers have committed vulnerable code. It evaluates 12 popular coding‑agent settings on these tasks and finds that all agents perform poorly in terms of security, with only 11.8% of solutions from SWE‑Agent with Claude 4 Sonnet being secure despite 57% being functionally correct. Attempts to mitigate security issues by adding vulnerability hints to the prompts do not improve results.

By Songwen Zhao, Danqing Wang, Kexun Zhang, Jiaxuan Luo, Zhuo Li, Lei Li
arXiv AI
Jun 2

Vibe-driven model-based engineering

arXiv:2604. 10645v2 Announce Type: replace-cross Abstract: There is a pressing need for better development methods and tools to keep up with the growing demand and increasing complexity of new software systems.

By Jordi Cabot
arXiv AI
Aug 24

Vibe Coding and Web Application Security: A Twin-Prompt Study

The study examines how adding a security-requirements section to prompts affects web applications generated by a large language model. Six distinct applications were produced twice—once with a baseline prompt and once with a security-aware prompt—yielding 12 programs. Analysis of these programs revealed 75 confirmed security findings, with the security-aware variants showing fewer issues (24 vs. 51) and no Critical or High severity problems.

By Darko Andro\v{c}ec