arXiv AI By Kirill Vasilevski (Justina), Ximing Dong (Justina), Benjamin Rombaut (Justina), Ruochen Deng (Justina), Jiahuei Lin (Justina), Arthur Leung, Dayi Lin, Boyuan Chen, Shaowei Wang, Ahmed E. Hassan

Beyond Correctness: Enhancing Architectural Reasoning in Code LLMs via Scalable Labeling with Agentic Judgment

Read the original on arXiv AI →

arXiv:2606. 14948v1 Announce Type: cross Abstract: LLMs have substantially improved software engineering yet real-world development requires architectural understanding.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
3d ago

E2E-SWE: Benchmarking LLMs on Building Working Codebases from Scratch

E2E-SWE is a benchmark that tests large language models’ ability to create complete, functional software repositories from scratch. It includes 186 tasks across 11 programming languages, each requiring an agent to build an installable project based solely on a natural‑language specification and an empty workspace, while passing a hidden test suite. The benchmark was crafted by software engineers and LLMs, then refined through iterative verification by autonomous agents to ensure clarity and solvability.

By Hantian Ding, Chloe Bi, Jiacheng Zhu, John Yang, Matt Deitke, Pengcheng Yin, Zijian Wang, Rui Hou
arXiv AI
Jun 2

Bridging Requirements and Architecture: Multi-Agent Orchestration with External Knowledge and Hierarchical Memory

arXiv:2606. 01385v1 Announce Type: cross Abstract: Software architecture design is a critical yet inherently complex and knowledge-intensive phase that requires balancing competing quality attributes and adapting to evolving requirements.

By Ruiyin Li, Yiran Zhang, Xiyu Zhou, Yangxiao Cai, Peng Liang, Weisong Sun, Jifeng Xuan, Zhi Jin, Yang Liu
arXiv AI
Jul 14

SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks

arXiv:2507. 11059v3 Announce Type: replace-cross Abstract: The rapid advancement of Large Language Models (LLMs) in software engineering has revealed critical limitations in existing benchmarks, particularly the widely used SWE-bench dataset.

By Pavel Adamenko, Mikhail Ivanov, Aidar Valeev, Rodion Levichev, Pavel Zadorozhny, Ivan Lopatin, Dmitry Babaev, Alena Fenogenova, Valentin Malykh
arXiv AI
Sep 4

Can LLMs Extract Architectural Design Decisions from Source Code Commits? - A Preliminary Exploratory Study

The study investigates whether large language models can extract Architectural Design Decisions (ADDs) from source code commits. Using four LLMs (Gemini 3 Pro, DeepSeek R1, Kimi K2, Qwen3) with zero‑shot and few‑shot prompting on 30 developer‑written ADDs, the authors evaluate outputs with ROUGE‑L, BLEU, METEOR, and BERTScore. Results show all models achieve a BERT‑F1 above 0.81, with few‑shot prompting slightly improving alignment, but the generated ADDs tend to be overly long, implementation‑focused, and lack the rationale behind the decisions.

By Amey Karan, Rudra Dhar, Mohamed Soliman, Karthik Vaidhyanathan