VibeMemBench: Evaluating Memory Systems for Coding Agents on Real Repository Coding Tasks
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
Schr"odinger's Repository (Schr"odingerRepo) is an evaluation framework that tests coding agents on dynamically instantiated repository representations to mitigate data leakage from static repository benchmarks. It transforms test repositories through four levels—problem statement reconstruction, namespace remapping, intra-file layout reordering, and functionality-preserving code rewriting—to obscure familiar cues while preserving executable behavior. Experiments on popular LLMs using SWE-bench Verified and SWE-QA show that removing these cues consistently degrades performance and increases interaction costs, mainly due to harder repository exploration and localization.
arXiv:2607. 28591v1 Announce Type: cross Abstract: Scaling coding agents requires a continuing supply of executable data for training, benchmarking, and continuous evaluation.
AdaRepair-Mem introduces an adaptive experience retrieval framework for repository-level program repair, addressing three key limitations in existing memory-augmented methods: imbalanced episodic memory, non-monotonic success with increased memory, and phase misalignment of memory types. The framework employs coverage-aware retrieval, quality-aware selection, and stage-aware routing to better match repair contexts with relevant memories. Evaluations on SWE-Bench-Lite and SWE-Bench-Verified show improved performance on under-covered repositories, reduced noisy retrieval, and enhanced patch refinement support.
arXiv:2606. 05646v1 Announce Type: cross Abstract: Large language models (LLMs) have enabled powerful software engineering (SE) agents capable of navigating complex codebases and resolving real-world issues.
arXiv:2607. 14390v1 Announce Type: cross Abstract: Coding agents now produce a growing share of a team's code, while the reasoning behind each change -- the alternatives weighed, the constraints discovered, the approaches rejected -- is trapped in assistant transcripts that vanish with the session.
Scaling coding agents requires a continuing supply of executable data for training, benchmarking, and continuous evaluation. Each task must couple a realistic software state with a specification, development tools, and reliable verification.