arXiv AI

AI-Augmented Closed-Loop Quality Engineering: A Reference Architecture for Continuous Software Quality Intelligence

arXiv:2606. 08793v1 Announce Type: cross Abstract: The quality of software engineering is still under a challenge due to disjointed processes between requirements, testing, and production, which hinders the opportunity to implement quality strategies in consecutive releases.

arXiv AI
Sep 1

Understanding Automated Program Repair Agents Through the Lens of Traceability: An Empirical Study

The paper presents a systematic analysis of five state‑of‑the‑art automated program repair agents, tracing their decision‑making across 500 real‑world repair tasks. It finds that while the agents perform well on simple fixes, they struggle with logic‑intensive bugs, often producing verbose, overfitted patches that pass tests without addressing root causes. Key bottlenecks identified include poor test generation, limited regression test selection, and reliance on primitive tooling without access to debuggers or advanced program analysis tools.

By Ira Ceka, Hailie Mitchell, Saurabh Pujar, Luca Buratti, Shyam Ramji, Junfeng Yang, Gail Kaiser, Baishakhi Ray
arXiv AI
Sep 10

Quality Metrics for LLM-Generated Asset Administration Shells: A Perturbation-Based Evaluation Approach

The paper introduces a perturbation-based framework to evaluate quality metrics for large language model (LLM) generated Asset Administration Shells (AAS). By systematically degrading AAS outputs across multiple dimensions, the study identifies that exact property name matching and value-based recall, along with name-based F1 score, best reflect quality changes. Experiments on 6,400 AAS instances from 200 products across GPT‑4o‑mini, Qwen3, and DeepSeek‑R1 reveal how different perturbations affect metrics and highlight variations among model families and product segments.

By Janek Gro{\ss}, Elena Zentgraf, Jens Heidrich
arXiv AI
Aug 28

From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench

The paper introduces MCR-Bench, a benchmark for realistic multi‑round code review that includes 2,269 real‑world tasks across five programming languages, each annotated with fine‑grained defect information and dynamic state labels. Experiments with mainstream large language models show limited overall performance, especially as interaction rounds increase, and reveal that model accuracy varies by defect type and severity. Error analysis identifies key failure mechanisms such as cross‑round temporal misalignment and insufficient long‑range memory.

By Dewu Zheng, Yanlin Wang, Xiwen Wang, Kefeng Duan, Hongyu Zhang, Xilin Liu, Yuchi Ma, Zibin Zheng
arXiv Machine Learning
Sep 14

SAGE-Loop: Reliable Closed-Loop LLM-Driven AutoML with Trial-and-Correction and Adaptive Ensembling

SAGE-Loop is a new closed‑loop, self‑adaptive AutoML framework that uses large language models to generate and validate machine learning pipelines in multiple rounds, allowing trial‑and‑repair and adaptive ensemble selection for both supervised and unsupervised tasks. It addresses the lack of instant feedback and correction in existing AutoML by enabling process‑level recovery from failures and dynamic use of model diversity. Experiments on 20 public datasets show consistent improvements in performance and stability across classification, regression, and clustering, and demonstrate the system’s ability to recover from execution failures.

By Junquan Gu, Shibo Cui, Xiangfeng Luo, Hang Yu
arXiv Computation and Language
Sep 2

Beneath the Diff: Diagnosing and Mitigating Algorithmic Mode Collapse in Code-Level Autonomous Research Loops

The paper investigates code-level autonomous research loops (ARLs) where a language model edits training pipelines to improve an in-loop metric. It identifies a failure mode called algorithmic mode collapse, where edits become semantically uniform despite surface diversity, leading to a growing gap between in-loop gains and independent evaluation. The authors propose Diversity‑Aware Proposal Sampling (DAPS), a lightweight method that reduces semantic decay by 69.1% and boosts faithfulness by over 80% while maintaining optimization speed.

By Bowei He, Weixu Zhang, Yili Jin, Xue Liu
arXiv AI
Sep 17

A Study of the Reliability of Agentic AI-Generated Programs

The paper investigates the reliability of software produced by agentic AI by comparing AI-generated versions of ten well-known Linux utilities to their human-written counterparts. Using fuzz testing (both black-box and coverage-guided AFL++), the authors find that AI-generated code is often as reliable or more reliable than the latest human versions, with fewer memory errors but a higher incidence of hangs. The study emphasizes that robust AI-generated software requires careful prompting, skilled human oversight, and that the AI workflow can serve as a cost-effective specification for sustainable code.

By Ayesha Shafique, Barton P. MIller, Elisa R. Heymann