arXiv AI

Can LLMs Extract Architectural Design Decisions from Source Code Commits? - A Preliminary Exploratory Study

The study investigates whether large language models can extract Architectural Design Decisions (ADDs) from source code commits. Using four LLMs (Gemini 3 Pro, DeepSeek R1, Kimi K2, Qwen3) with zero‑shot and few‑shot prompting on 30 developer‑written ADDs, the authors evaluate outputs with ROUGE‑L, BLEU, METEOR, and BERTScore. Results show all models achieve a BERT‑F1 above 0.81, with few‑shot prompting slightly improving alignment, but the generated ADDs tend to be overly long, implementation‑focused, and lack the rationale behind the decisions.

arXiv AI
Jun 16

Beyond Correctness: Enhancing Architectural Reasoning in Code LLMs via Scalable Labeling with Agentic Judgment

arXiv:2606. 14948v1 Announce Type: cross Abstract: LLMs have substantially improved software engineering yet real-world development requires architectural understanding.

By Kirill Vasilevski (Justina), Ximing Dong (Justina), Benjamin Rombaut (Justina), Ruochen Deng (Justina), Jiahuei Lin (Justina), Arthur Leung, Dayi Lin, Boyuan Chen, Shaowei Wang, Ahmed E. Hassan
arXiv AI
Aug 19

GADR: Gathering Architecture Decision Records from Meeting Transcriptions

The paper introduces GADR, a multi‑agent, self‑correcting workflow that extracts architectural decisions from raw meeting transcriptions and produces Nygard‑formatted ADR drafts. A feasibility study using five real project transcripts, expert reviews by four senior architects, and evaluations by fifteen students shows that GADR captures most expert‑identified decisions and yields drafts that participants find clear and useful, outperforming zero‑shot and few‑shot baselines in stability and structural adherence. The study also highlights a trade‑off: RAG‑based enrichment can deepen ADR content but may introduce transcript‑unfaithful material, raising open questions about traceability in automated architectural documentation.

By Lucas Daniel Costa da Silva, Kiev Gama
arXiv AI
Sep 18

BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software

BuildBench introduces a realistic benchmark for evaluating large language model agents on the task of compiling open‑source software (OSS). It includes diverse OSS projects that lack clear build instructions, have undocumented dependencies, and may require source patching or script modification. The authors also present OSS‑BUILD‑AGENT, a baseline LLM‑based agent that retrieves build instructions effectively and achieves state‑of‑the‑art performance on the benchmark.

By Zehua Zhang, Ati Priya Bajaj, Divij Handa, Siyu Liu, Arvind S Raj, Hongkai Chen, Hulin Wang, Yibo Liu, Zion Leonahenahe Basque, Souradip Nath, Vishal Juneja, Nikhil Chapre, Tiffany Bao, Yan Shoshitaishvili, Adam Doup\'e, Chitta Baral, Ruoyu Wang
arXiv AI
Jun 12

HalluJudge: A Reference-Free Hallucination Detection for Context Misalignment in Code Review Automation

arXiv:2601. 19072v3 Announce Type: replace-cross Abstract: Large Language models (LLMs) have shown strong capabilities in code review automation, such as review comment generation, yet they suffer from hallucinations -- where the generated review comments are ungrounded in the actual code -- poses a significant challenge to the adoption of LLMs in code review workflows.

By Kla Tantithamthavorn, Hong Yi Lin, Patanamon Thongtanunam, Wachiraphan Charoenwet, Minwoo Jeong, Ming Wu