arXiv:2606. 18293v1 Announce Type: cross Abstract: Thanks to rapid developments in generative AI, we are in the midst of a paradigm shift that may change how we interact with computers forever.
By Callum Barbour
The paper introduces SUSVIBES, a benchmark of 186 real‑world software engineering tasks where human programmers have committed vulnerable code. It evaluates 12 popular coding‑agent settings on these tasks and finds that all agents perform poorly in terms of security, with only 11.8% of solutions from SWE‑Agent with Claude 4 Sonnet being secure despite 57% being functionally correct. Attempts to mitigate security issues by adding vulnerability hints to the prompts do not improve results.
By Songwen Zhao, Danqing Wang, Kexun Zhang, Jiaxuan Luo, Zhuo Li, Lei Li
The unified agent for long-horizon productivity and coding, launching with Work and Code modes. Plus, a new Vibe VS Code extension.
arXiv:2604. 14137v3 Announce Type: replace-cross Abstract: Evaluating LLMs is challenging, as benchmark scores often fail to capture models' real-world usefulness.
By Itay Itzhak, Eliya Habba, Gabriel Stanovsky, Yonatan Belinkov
The study evaluates Vibe Coding, an AI‑led conversational programming paradigm that lets developers generate software via natural‑language interaction with large language models. In a mixed‑methods experiment with 30 participants, Vibe Coding improved development efficiency—reducing task completion time by 27% versus traditional coding and 12% versus AI‑assisted coding—while also yielding a good usability score (SUS = 71.4) and moderate cognitive workload (NASA‑TLX = 55.5). However, the gains came with trade‑offs: lower maintainability indices, higher security vulnerabilities, and themes of trust calibration, loss of control, and prompt‑engineering strategy emerged, leading the authors to propose a three‑pillar framework for responsible adoption.
By Sales G. Aribe Jr., Louie Jay S. Labastida
The paper introduces a new inference architecture for vibe design agents that separates design exploration from implementation. By generating structured design specifications with typicality scores and selecting one for downstream generation, the method allows users to explore coherent UI alternatives without altering the underlying generation settings. Experiments on UI themes and visual-asset prompts show increased selection coverage and screenshot variation, with mixed preferences from an LLM judge and modest operational costs in a large online test.
By Yifan Zhang, Nghi D. Q. Bui, Georgios Evangelopoulos, Arnaud Benard