arXiv:2609.16936v1 Announce Type: cross
Abstract: Large language model (LLM)-powered coding agents have made rapid progress in automating software engineering tasks, yet repository-level issue resolu...
By Yunxiang Zhang, Haiquan Wang, JiaWei Guo, Hanyang Xia, Yan Chen, Tong Chen, Zhang Zhiwei, Junchen Ye
PPTBench is a new benchmark that tests coding agents’ ability to reconstruct scientific flow diagrams from arXiv papers into editable PowerPoint slides. The dataset contains 500 tasks, each requiring agents to produce a single PPTX page with native, editable objects, and a four‑stage Agentic Judge evaluates validity, semantic correctness, rendering quality, and fine‑grained visual quality. Across 31 model configurations, the best score is 67.80, with a median of 19.47, showing that while agents can generate valid PPTX files, they still struggle with semantic and visual accuracy, especially text details.
By Xiaoqiu Wang, Yizhe Chi, Wenyi Li, Deyao Hong, Zhihan Shan, Mingju Gao, Kaisen Yang, Youjie Zheng, Calvin Xiao, Qinhuai Na
The paper explores whether natural‑language documentation aids coding agents in fixing software bugs and introduces a roundtrip benchmark that evaluates code descriptions by regenerating code and testing it. It finds that description completeness, not length, determines fidelity, and presents an optimizer that can produce fully faithful descriptions that generalize to new files. However, experiments across two model families and ten repositories show that such compact documentation does not improve an agent’s ability to resolve real repository issues compared to using the issue alone.
By Md Shohel Arman, Igor Molybog
arXiv:2606.21804v2 Announce Type: replace-cross
Abstract: Maintainability is a core dimension of software engineering, shaping how code is written, reviewed, and developed over time. While coding age...
By Shaswat Patel, Betty Li Hou, Arun Purohit, Kai Xu, Jane Pan, He He, Valerie Chen
arXiv:2602.11988v3 Announce Type: replace-cross
Abstract: A widespread practice in software development is to tailor coding agents to repositories using context files, such as AGENTS.md. Although thi...
By Thibaud Gloaguen, Niels M\"undler-Sasahara, Mark Niklas M\"uller, Veselin Raychev, Martin Vechev
arXiv:2605. 26144v2 Announce Type: replace-cross Abstract: We present VISTA (VIsual Spec-To-App Benchmark), a benchmark for evaluating the end-to-end web-app generation capabilities of LLM-based agents.
By JunJia Guo (Joe), Yuhang Yao (Joe), Jiawei (Joe), Zhou, Jingdi Chen