arXiv AI

Vibe Coding Ate My Homework: An evaluation of AI approaches to greenfield software engineering and programming

arXiv:2606. 18293v1 Announce Type: cross Abstract: Thanks to rapid developments in generative AI, we are in the midst of a paradigm shift that may change how we interact with computers forever.

arXiv AI
Jun 2

Vibe-driven model-based engineering

arXiv:2604. 10645v2 Announce Type: replace-cross Abstract: There is a pressing need for better development methods and tools to keep up with the growing demand and increasing complexity of new software systems.

By Jordi Cabot
arXiv AI
Sep 11

The Vibe Shift in Software Engineering: Evaluating AI-Led Conversational Programming for Performance, Cognition, and Responsible Adoption

The study evaluates Vibe Coding, an AI‑led conversational programming paradigm that lets developers generate software via natural‑language interaction with large language models. In a mixed‑methods experiment with 30 participants, Vibe Coding improved development efficiency—reducing task completion time by 27% versus traditional coding and 12% versus AI‑assisted coding—while also yielding a good usability score (SUS = 71.4) and moderate cognitive workload (NASA‑TLX = 55.5). However, the gains came with trade‑offs: lower maintainability indices, higher security vulnerabilities, and themes of trust calibration, loss of control, and prompt‑engineering strategy emerged, leading the authors to propose a three‑pillar framework for responsible adoption.

By Sales G. Aribe Jr., Louie Jay S. Labastida
arXiv Computation and Language
Aug 24

Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks

The paper introduces SUSVIBES, a benchmark of 186 real‑world software engineering tasks where human programmers have committed vulnerable code. It evaluates 12 popular coding‑agent settings on these tasks and finds that all agents perform poorly in terms of security, with only 11.8% of solutions from SWE‑Agent with Claude 4 Sonnet being secure despite 57% being functionally correct. Attempts to mitigate security issues by adding vulnerability hints to the prompts do not improve results.

By Songwen Zhao, Danqing Wang, Kexun Zhang, Jiaxuan Luo, Zhuo Li, Lei Li
arXiv AI
Jun 9

Lost in the Flow with Code Talkers: Unveiling the Instruction-Tuning Tax of Large Language Models in Code Tasks

arXiv:2606. 08676v1 Announce Type: cross Abstract: AI coding assistants have significantly improved developer productivity by automatically suggesting code that aligns with user intent, and many of these tools are now integrated directly into Integrated Development Environments (IDEs).

By Shi Ying Chang, Chiok Yew Ho, Yichen Li, Yintong Huo
arXiv AI
Aug 24

Vibe Coding and Web Application Security: A Twin-Prompt Study

The study examines how adding a security-requirements section to prompts affects web applications generated by a large language model. Six distinct applications were produced twice—once with a baseline prompt and once with a security-aware prompt—yielding 12 programs. Analysis of these programs revealed 75 confirmed security findings, with the security-aware variants showing fewer issues (24 vs. 51) and no Critical or High severity problems.

By Darko Andro\v{c}ec