arXiv Computation and Language

VibeMemBench: Evaluating Memory Systems for Coding Agents on Real Repository Coding Tasks

arXiv AI
Sep 24

Schr\"odinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It?

Schr"odinger's Repository (Schr"odingerRepo) is an evaluation framework that tests coding agents on dynamically instantiated repository representations to mitigate data leakage from static repository benchmarks. It transforms test repositories through four levels—problem statement reconstruction, namespace remapping, intra-file layout reordering, and functionality-preserving code rewriting—to obscure familiar cues while preserving executable behavior. Experiments on popular LLMs using SWE-bench Verified and SWE-QA show that removing these cues consistently degrades performance and increases interaction costs, mainly due to harder repository exploration and localization.

By Silin Chen, Yufei Yang, Xiaodong Gu, Yuling Shi, Chengcheng Wan, Haibing Guan
arXiv AI
Sep 18

AdaRepair-Mem: Adaptive Experience Orchestration for Repository-Level Program Repair

AdaRepair-Mem introduces an adaptive experience retrieval framework for repository-level program repair, addressing three key limitations in existing memory-augmented methods: imbalanced episodic memory, non-monotonic success with increased memory, and phase misalignment of memory types. The framework employs coverage-aware retrieval, quality-aware selection, and stage-aware routing to better match repair contexts with relevant memories. Evaluations on SWE-Bench-Lite and SWE-Bench-Verified show improved performance on under-covered repositories, reduced noisy retrieval, and enhanced patch refinement support.

By Z. C. Luo, J. C. Guo, W. J. He, S. Y. Wang, J. C. Yu, F. M. Zhao, Y. Chen, T. Cao, L. Q. Liu, N. Zheng, W. Xu, J. Jiang, Z. M. Zhao
arXiv AI
3d ago

Zero2Repo: Can Coding Agents Build Repositories from Scratch?

arXiv:2609.38269v1 Announce Type: cross Abstract: Coding agents are increasingly asked to build software rather than patch it, yet benchmarks for from-scratch repository construction are mostly limit...

By Pei Yang, Tianyu Shi, Yuhang Yao, Wanyi Chen, Tongyun Yang, Dun Pei, Haonan Wang, Pengbin Feng, Guanxu Yu, Jingchun Huang, Zeyu Zhang, Shuhan Sun, Hao Li, Xiang Li, Jie Xiao, Xinyu Wang, Hanxin Chen, Daqi Li, Qi Jia, Hongshan Lin, Zhizhou Gu, Zijun Tian, Weizhi Du, Lynn Ai, Eric Yang
arXiv AI
Sep 10

When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents

The paper introduces MERIT, a benchmark that evaluates the marginal benefit of long‑term memory for tool‑using large language model agents while explicitly accounting for cost. MERIT provides episodic tool‑use tasks across three domains, verifies dependence on earlier‑episode facts, and measures memory operations in tokens and dollars. Experiments on GPT‑4.1‑mini, Claude Haiku 4.5, and Claude Sonnet 5 show that memory can significantly improve task success, but its utility varies widely across models and memory implementations, and full replay is rarely cost‑effective.

By Shweta Mishra, Shashank Mishra
arXiv AI
Aug 11

A Unified Issue Resolution Benchmark for Requirement Clarification, Planning, and Code Generation for Coding Agents

arXiv:2608. 09072v1 Announce Type: cross Abstract: Large language model-powered coding agents are increasingly used to modify existing code repositories, for example, by adding features or fixing bugs.

By Xin Zhou, Chun Yong Chong, Kisub Kim, Yun Peng, Rui Shu, Zihan Wu, Xu Han, Guowen Yuan, Zeyang Zhuang, Jounghoon Kim, Jeongjin Ju, Seongmin Ju, Taein Yoon, David Lo