arXiv AI

Information-Gain Rewards over Diversity-Pruned Tests: GT-Anchored Verifier Co-Training for Reliable Code Generation

arXiv Machine Learning
Jun 26

Reinforcement Learning without Ground-Truth Solutions can Improve LLMs

arXiv:2606. 27369v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) for training LLMs typically rely on ground-truth answers to assign rewards, limiting their applicability to tasks where the ground-truth solution is unknown.

By Yingyu Lin, Qiyue Gao, Nikki Lijing Kuang, Xunpeng Huang, Kun Zhou, Tongtong Liang, Zhewei Yao, Yi-An Ma, Yuxiong He
arXiv AI
Aug 13

Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL

arXiv:2608. 11669v1 Announce Type: cross Abstract: Reinforcement learning against rubrics, lists of criteria graded by an LLM judge, has become a standard way to post-train language models on tasks with no deterministic answer.

By Minglai Yang, Xinyu Guo, Utkarsh Tyagi, Mian Zhang, Razvan Dumitru, Sunjie Hou, Yunzhong He, Daniel Yue Zhang, Ying Liu
arXiv Machine Learning
Sep 1

X-Coder: Advancing Competitive Programming with Synthetic Tasks, Solutions, and Tests

The paper introduces X-Coder, a competitive programming model trained entirely on synthetic tasks, verified solutions, and reliable test cases, eliminating the need for real-world data in post‑training. A dual‑verification strategy is used to reduce noise in solutions and test outputs, providing high‑quality reward signals for reinforcement learning. X‑Coder‑14B achieves significant performance gains, scoring 67.5% on LiveCodeBench v5 and 63.4% on v6, surpassing its base model by over 40 points.

By Jie Wu, Haoling Li, Xin Zhang, Jiani Guo, Jane Luo, Xuewei Yang, Steven Liu, Yangyu Huang, Ruihang Chu, Scarlett Li, Yujiu Yang
arXiv AI
Sep 7

What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents

The study investigates how multi‑harness reinforcement learning (RL) affects coding agents by comparing two grouping strategies—Within (one group per task‑harness pair) and Cross (harnesses pooled within a task)—using a Qwen3‑8B policy trained on frozen task‑harness records from Aider, OpenHands, Qwen Code, and SWE‑agent. Across 24,000 sealed evaluations, the choice of evaluation harness dramatically increases solve rates (from 2.14 % to 9.27 %), while the grouping rule has a negligible effect. Both grouping rules yield similar gains on the same source harness, and Cross‑harness credit does not improve portability beyond Within‑harness credit, suggesting that multi‑harness RL reports should specify grouping boundaries and test on unseen harnesses.

By Chenqian Le, Jiayi Cheng, Qijia He, Runhao Li, Yinghao Li, Xupeng Chen
arXiv AI
Jul 16

Representation-Based Exploration for Language Models: From Test-Time to Post-Training

arXiv:2510. 11686v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) promises to expand the capabilities of language models, but it is unclear if current RL techniques promote the discovery of novel behaviors, or simply sharpen those already present in the base model.

By Jens Tuyls, Dylan J. Foster, Akshay Krishnamurthy, Jordan T. Ash
arXiv AI
Sep 4

ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize

ESPO (Error-Structured Prompt Optimization) addresses prompt bloat in evolutionary prompt optimizers by splitting the optimization process into Diagnose, Propose, and Select phases. It clusters training errors into structural patterns, generates diverse candidate prompts through four complementary strategies, and applies bootstrap stability selection. Across seven NLP benchmarks, ESPO improves average accuracy by +3.76 pp over GEPA, produces prompts 47 % shorter, and achieves higher accuracy on four additional student models, with the largest gain on Qwen3 GSM8K.

By Lihao Liu, Peng Tang, Kunwar Yashraj Singh, Shabnam Ghadar
arXiv Machine Learning
Sep 1

Agnostics: Learning to Code in Any Programming Language via Reinforcement with a Universal Learning Environment

Agnostics is a language‑agnostic post‑training pipeline that uses reinforcement learning with verifiable rewards (RLVR) to improve large language models on low‑resource programming languages. By rewriting unit‑test datasets into a language‑independent I/O format, providing a short configuration for compiling and running code, and employing a single verifier that judges code by observable behavior, Agnostics eliminates the need for language‑specific engineering. Applied to Lua, Julia, R, OCaml, and Fortran, it boosts Qwen‑3 4B to rival larger models, scales to diverse families, and achieves new state‑of‑the‑art pass@1 on MultiPL‑E and a new multi‑language LiveCodeBench.

By Aleksander Boruch-Gruszecki, Yangtian Zi, Zixuan Wu, Tejas Oberoi, Carolyn Jane Anderson, Joydeep Biswas, Arjun Guha