arXiv AI By Amey Karan, Rudra Dhar, Mohamed Soliman, Karthik Vaidhyanathan

Can LLMs Extract Architectural Design Decisions from Source Code Commits? - A Preliminary Exploratory Study

Read the original on arXiv AI →

The study investigates whether large language models can extract Architectural Design Decisions (ADDs) from source code commits. Using four LLMs (Gemini 3 Pro, DeepSeek R1, Kimi K2, Qwen3) with zero‑shot and few‑shot prompting on 30 developer‑written ADDs, the authors evaluate outputs with ROUGE‑L, BLEU, METEOR, and BERTScore. Results show all models achieve a BERT‑F1 above 0.81, with few‑shot prompting slightly improving alignment, but the generated ADDs tend to be overly long, implementation‑focused, and lack the rationale behind the decisions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 16

Beyond Correctness: Enhancing Architectural Reasoning in Code LLMs via Scalable Labeling with Agentic Judgment

arXiv:2606. 14948v1 Announce Type: cross Abstract: LLMs have substantially improved software engineering yet real-world development requires architectural understanding.

By Kirill Vasilevski (Justina), Ximing Dong (Justina), Benjamin Rombaut (Justina), Ruochen Deng (Justina), Jiahuei Lin (Justina), Arthur Leung, Dayi Lin, Boyuan Chen, Shaowei Wang, Ahmed E. Hassan
arXiv AI
Aug 19

GADR: Gathering Architecture Decision Records from Meeting Transcriptions

The paper introduces GADR, a multi‑agent, self‑correcting workflow that extracts architectural decisions from raw meeting transcriptions and produces Nygard‑formatted ADR drafts. A feasibility study using five real project transcripts, expert reviews by four senior architects, and evaluations by fifteen students shows that GADR captures most expert‑identified decisions and yields drafts that participants find clear and useful, outperforming zero‑shot and few‑shot baselines in stability and structural adherence. The study also highlights a trade‑off: RAG‑based enrichment can deepen ADR content but may introduce transcript‑unfaithful material, raising open questions about traceability in automated architectural documentation.

By Lucas Daniel Costa da Silva, Kiev Gama