A^2Agent: Action-Aware Reinforcement Learning for Repository-Level Code Localization Agents
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
arXiv:2608.29675v1 Announce Type: cross Abstract: Repository exploration is a distinct and costly stage of coding-agent pipelines: before generating a patch, an agent must identify which repository f...
BoostAPR is a three-stage framework that improves automated program repair by using execution-grounded reinforcement learning with dual reward models. The approach first fine‑tunes a model on execution‑verified demonstrations, then trains a sequence‑level assessor and a line‑level credit allocator from execution outcomes, and finally applies PPO optimization where the line‑level model redistributes rewards to critical edit regions. Evaluated on SWE‑Gym and four benchmarks, BoostAPR achieves significant gains, including 40.7% on SWE‑bench Verified and 95.0% on QuixBugs, demonstrating strong cross‑language generalization.
arXiv:2608.27906v2 Announce Type: replace Abstract: Interactive web application generation requires models to produce usable HTML, CSS, and JavaScript applications from natural language requests. Unl...
arXiv:2609.22068v1 Announce Type: new Abstract: Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers. Open-source codebases offer a rich sourc...
LEGO-RL is a framework that connects native coding-agent harnesses with scalable policy‑gradient training without altering the harnesses’ internal flow. It achieves faithful optimization through in‑process LLM proxying, reliable execution via sandbox orchestration, and observable training with automated validation and a Live UI. Experiments show LEGO‑RL improves the Qwen3.5‑35B‑A3B model’s performance on three native harnesses while preserving high rollout‑training probability correlation.
arXiv:2607. 08837v1 Announce Type: cross Abstract: Exploration is essential to RL since a policy cannot improve by repeatedly sampling the behaviors it already prefers.