arXiv AI

BoostAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward Models

BoostAPR is a three-stage framework that improves automated program repair by using execution-grounded reinforcement learning with dual reward models. The approach first fine‑tunes a model on execution‑verified demonstrations, then trains a sequence‑level assessor and a line‑level credit allocator from execution outcomes, and finally applies PPO optimization where the line‑level model redistributes rewards to critical edit regions. Evaluated on SWE‑Gym and four benchmarks, BoostAPR achieves significant gains, including 40.7% on SWE‑bench Verified and 95.0% on QuixBugs, demonstrating strong cross‑language generalization.

arXiv AI
6d ago

SLMFix: Leveraging Small Language Models for Domain Specific Language Error Fixing with Reinforcement Learning

SLMFix is a code‑generation pipeline that uses a small language model fine‑tuned with reinforcement learning to correct syntactic errors in programs produced by large language models for domain‑specific languages. The approach relies on interpreter feedback to guide the error‑fixing process. Experiments show that SLMFix improves validator pass rates by 40% on low‑resource programming languages and removes over 50% of syntactic errors on high‑resource DSLs, outperforming supervised fine‑tuning even for 7B models.

By David Jiahao Fu, Aryan Gupta, Aaron Councilman, Yu-Xiong Wang, Vikram Adve
arXiv AI
Aug 18

ClawGym II: Exploring Black-Box RL on Agent Harness

arXiv:2608. 16798v1 Announce Type: cross Abstract: Agent harnesses have substantially improved performance on long-horizon tasks by coordinating agent interactions with the environment.

By Huatong Song, Fei Bai, Ming Yang, Renyuan Li, Jia Deng, Jujie He, Zhange Zhang, Daixuan Cheng, Yan Xing, Qi Yun, Xuxing Chen, Danyang Li, Feng Chang, Chuan Hao, Ran Tao, Jian Yang, Bryan Dai, Wayne Xin Zhao, Mingjie Tang, Ji-Rong Wen