The study investigates whether large language models can extract Architectural Design Decisions (ADDs) from source code commits. Using four LLMs (Gemini 3 Pro, DeepSeek R1, Kimi K2, Qwen3) with zero‑shot and few‑shot prompting on 30 developer‑written ADDs, the authors evaluate outputs with ROUGE‑L, BLEU, METEOR, and BERTScore. Results show all models achieve a BERT‑F1 above 0.81, with few‑shot prompting slightly improving alignment, but the generated ADDs tend to be overly long, implementation‑focused, and lack the rationale behind the decisions.
By Amey Karan, Rudra Dhar, Mohamed Soliman, Karthik Vaidhyanathan
The study investigates how users interact with ChatGPT for code generation beyond simple function-level tasks, focusing on project-level benchmarks that involve multi-class dependencies. A user study with 36 participants examined prompting patterns, screen recordings, and chat logs to identify Human‑LLM Interaction (HLI) features that influence productivity. The results highlight three consistently supportive HLI features, five guidelines to boost productivity, and a taxonomy of 29 runtime and logic errors with mitigation strategies.
By Sangwon Hyun, Hyunjun Kim, Jinhyuk Jang, Hyojin Choi, M. Ali Babar
arXiv:2609.22102v1 Announce Type: cross
Abstract: Large Language Models (LLMs) are widely used as programming assistants, yet it remains unclear whether and how user's demographic information impacts...
By Anubhav Gupta, Mayara Costa Figueiredo, Leticia Santos Machado, Tanner Wright, Ivan Beschastnikh, Cleidson R. B. de Souza, Gema Rodr\'iguez-P\'erez
The use of LLMs in software development has become increasingly widespread on tasks such as code generation and summarization. Reports from large technology companies showed that around 20% to 30% of their code are generated by LLMs.
arXiv:2607. 01867v1 Announce Type: cross Abstract: The use of LLMs in software development has become increasingly widespread on tasks such as code generation and summarization.
By Yongyi Ji, Jiaji Wang, Yi Zhou, Fuxiang Chen, Hongji Yang
arXiv:2608. 11513v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly integrated into software engineering workflows, helping developers write, debug, test, and maintain code.
By Alex Deaconu, Anubhav Gupta, Manaal Basha, Nicholas Haydu, Gema Rodr\'iguez-P\'erez
The study evaluates Vibe Coding, an AI‑led conversational programming paradigm that lets developers generate software via natural‑language interaction with large language models. In a mixed‑methods experiment with 30 participants, Vibe Coding improved development efficiency—reducing task completion time by 27% versus traditional coding and 12% versus AI‑assisted coding—while also yielding a good usability score (SUS = 71.4) and moderate cognitive workload (NASA‑TLX = 55.5). However, the gains came with trade‑offs: lower maintainability indices, higher security vulnerabilities, and themes of trust calibration, loss of control, and prompt‑engineering strategy emerged, leading the authors to propose a three‑pillar framework for responsible adoption.
By Sales G. Aribe Jr., Louie Jay S. Labastida
arXiv:2608.30756v1 Announce Type: cross
Abstract: Large language models (LLMs) have become an essential tool for assisting developers, yet we still lack knowledge on ways to effectively support their...
By Annemarie Wittig, Alina Mailach, Janet Siegmund, Norbert Siegmund
Context: Generative AI (GenAI) and Large Language Models (LLMs) are increasingly used for academic tasks in software engineering and beyond, including systematic literature reviews (SLRs). However, while capable of summarizing text, there is no guarantee they can meet the rigour, reliability, and transparency that SLRs require.
arXiv:2508. 16131v3 Announce Type: replace-cross Abstract: Code completion entails the task of providing missing tokens given a surrounding context.
By Zoe Kotti, Konstantina Dritsa, Diomidis Spinellis, Panos Louridas
arXiv:2601. 19072v3 Announce Type: replace-cross Abstract: Large Language models (LLMs) have shown strong capabilities in code review automation, such as review comment generation, yet they suffer from hallucinations -- where the generated review comments are ungrounded in the actual code -- poses a significant challenge to the adoption of LLMs in code review workflows.
By Kla Tantithamthavorn, Hong Yi Lin, Patanamon Thongtanunam, Wachiraphan Charoenwet, Minwoo Jeong, Ming Wu
arXiv:2606. 29520v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used as assistants across the software development lifecycle, yet their ability to reason about software architecture remains largely unmeasured.
By Tiziano Santilli, Francesco Daghero, Mayhar Tourchi Moghaddam