arXiv:2505. 07372v3 Announce Type: replace-cross Abstract: This paper presents a novel methodology for enhancing Automated Program Repair (APR) through synthetic data generation utilizing Large Language Models (LLMs).
By David de-Fitero-Dominguez, Antonio Garcia-Cabot, Eva Garcia-Lopez
The paper presents a systematic analysis of five state‑of‑the‑art automated program repair agents, tracing their decision‑making across 500 real‑world repair tasks. It finds that while the agents perform well on simple fixes, they struggle with logic‑intensive bugs, often producing verbose, overfitted patches that pass tests without addressing root causes. Key bottlenecks identified include poor test generation, limited regression test selection, and reliance on primitive tooling without access to debuggers or advanced program analysis tools.
By Ira Ceka, Hailie Mitchell, Saurabh Pujar, Luca Buratti, Shyam Ramji, Junfeng Yang, Gail Kaiser, Baishakhi Ray
The paper introduces a perturbation-based framework to evaluate quality metrics for large language model (LLM) generated Asset Administration Shells (AAS). By systematically degrading AAS outputs across multiple dimensions, the study identifies that exact property name matching and value-based recall, along with name-based F1 score, best reflect quality changes. Experiments on 6,400 AAS instances from 200 products across GPT‑4o‑mini, Qwen3, and DeepSeek‑R1 reveal how different perturbations affect metrics and highlight variations among model families and product segments.
By Janek Gro{\ss}, Elena Zentgraf, Jens Heidrich
arXiv:2606. 24968v1 Announce Type: new Abstract: Context: Software defect prediction supports maintenance decisions such as testing prioritization, release-risk assessment, and quality monitoring.
By Emmanuel Charleson Dapaah, Philip Makedonski, Jens Grabowski
The paper introduces MCR-Bench, a benchmark for realistic multi‑round code review that includes 2,269 real‑world tasks across five programming languages, each annotated with fine‑grained defect information and dynamic state labels. Experiments with mainstream large language models show limited overall performance, especially as interaction rounds increase, and reveal that model accuracy varies by defect type and severity. Error analysis identifies key failure mechanisms such as cross‑round temporal misalignment and insufficient long‑range memory.
By Dewu Zheng, Yanlin Wang, Xiwen Wang, Kefeng Duan, Hongyu Zhang, Xilin Liu, Yuchi Ma, Zibin Zheng
SAGE-Loop is a new closed‑loop, self‑adaptive AutoML framework that uses large language models to generate and validate machine learning pipelines in multiple rounds, allowing trial‑and‑repair and adaptive ensemble selection for both supervised and unsupervised tasks. It addresses the lack of instant feedback and correction in existing AutoML by enabling process‑level recovery from failures and dynamic use of model diversity. Experiments on 20 public datasets show consistent improvements in performance and stability across classification, regression, and clustering, and demonstrate the system’s ability to recover from execution failures.
By Junquan Gu, Shibo Cui, Xiangfeng Luo, Hang Yu