RLPF: Reinforcement Learning from Performance Feedback for Code Generation
arXiv:2607. 27271v1 Announce Type: new Abstract: Code models are increasingly trained with execution feedback, but most training signals still stop at correctness.
BoostAPR is a three-stage framework that improves automated program repair by using execution-grounded reinforcement learning with dual reward models. The approach first fine‑tunes a model on execution‑verified demonstrations, then trains a sequence‑level assessor and a line‑level credit allocator from execution outcomes, and finally applies PPO optimization where the line‑level model redistributes rewards to critical edit regions. Evaluated on SWE‑Gym and four benchmarks, BoostAPR achieves significant gains, including 40.7% on SWE‑bench Verified and 95.0% on QuixBugs, demonstrating strong cross‑language generalization.
arXiv:2607. 27271v1 Announce Type: new Abstract: Code models are increasingly trained with execution feedback, but most training signals still stop at correctness.
arXiv:2606. 08346v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a dominant paradigm for improving the reasoning capabilities of large language models (LLMs).
arXiv:2507. 22580v2 Announce Type: replace-cross Abstract: Automated Program Repair (APR) seeks to automatically correct software bugs without requiring human intervention.
arXiv:2608.27906v2 Announce Type: replace Abstract: Interactive web application generation requires models to produce usable HTML, CSS, and JavaScript applications from natural language requests. Unl...
arXiv:2603. 19329v3 Announce Type: replace-cross Abstract: Large language models (LLMs) can generate plausible code but offer limited guarantees of correctness.
arXiv:2606. 20068v1 Announce Type: new Abstract: While reinforcement learning from verifiable rewards (RLVR) typically has relied on a single binary verification signal, symbolic proof assistants in formal reasoning offer rich, fine-grained structured feedback.
SLMFix is a code‑generation pipeline that uses a small language model fine‑tuned with reinforcement learning to correct syntactic errors in programs produced by large language models for domain‑specific languages. The approach relies on interpreter feedback to guide the error‑fixing process. Experiments show that SLMFix improves validator pass rates by 40% on low‑resource programming languages and removes over 50% of syntactic errors on high‑resource DSLs, outperforming supervised fine‑tuning even for 7B models.
arXiv:2608. 16798v1 Announce Type: cross Abstract: Agent harnesses have substantially improved performance on long-horizon tasks by coordinating agent interactions with the environment.
arXiv:2608. 01804v1 Announce Type: new Abstract: Post-training large language models (LLMs) via reinforcement learning (RL) has significantly advanced code generation capabilities.
Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, making fine-grained credit assignment across multi-step interactions difficult. Existing approache...
arXiv:2604. 00860v3 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become a central post-training paradigm for improving the reasoning capabilities of large language models.
arXiv:2601. 03555v3 Announce Type: replace Abstract: Training reliable tool-augmented agents remains a significant challenge, largely due to the difficulty of credit assignment in multi-step reasoning.