arXiv AI

Cross-Benchmark Transfer from RL on Agentic Coding Tasks

The paper reports that applying reinforcement learning (RL) to the Kimi K2.7 Code model on 1,700 agentic coding tasks improves its performance on six external benchmarks. After a single epoch of GSPO training on a rank‑32 LoRA adapter, pass‑@1 scores increased across all benchmarks, with significant gains even on data released after training. The trained model also reduces agent steps and avoids common failure modes such as dropping requirements or breaking existing behavior.

arXiv AI
Sep 7

What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents

The study investigates how multi‑harness reinforcement learning (RL) affects coding agents by comparing two grouping strategies—Within (one group per task‑harness pair) and Cross (harnesses pooled within a task)—using a Qwen3‑8B policy trained on frozen task‑harness records from Aider, OpenHands, Qwen Code, and SWE‑agent. Across 24,000 sealed evaluations, the choice of evaluation harness dramatically increases solve rates (from 2.14 % to 9.27 %), while the grouping rule has a negligible effect. Both grouping rules yield similar gains on the same source harness, and Cross‑harness credit does not improve portability beyond Within‑harness credit, suggesting that multi‑harness RL reports should specify grouping boundaries and test on unseen harnesses.

By Chenqian Le, Jiayi Cheng, Qijia He, Runhao Li, Yinghao Li, Xupeng Chen
arXiv Machine Learning
Sep 11

T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks

The paper introduces T1, a 122‑billion‑parameter Mixture‑of‑Experts model trained with reinforcement learning to perform long‑horizon terminal tasks such as coding and scientific discovery. T1 operates a real shell in a cloud sandbox, making over 300 tool‑call turns per task and receiving rewards from task‑specific verifiers. The authors detail a training recipe that includes aggressive warm‑starting, TITO construction with drift repair, and rollout‑routing replay, achieving significant performance gains on Terminal‑Bench 2.1 and surpassing GPT‑5.4 and GLM‑5.1 on the Long‑Horizon Terminal Bench.

By Junyao Yang, Yucheng Shi, Zhongzhi Li, Ruhan Wang, Zongxia Li, Haitao Mi, Leowei Liang
arXiv AI
3d ago

Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

Mid‑Harness proposes a test‑time compute strategy that samples and verifies candidate actions before execution, keeping the underlying generator and harness unchanged. Experiments show that with a strong verifier, sampling more actions significantly boosts success rates—e.g., a GPT‑5.6 verifier raises Pass@1 from 50.00 % to 68.03 % on TerminalBench‑Lite using eight samples. The approach also improves performance across various models, benchmarks, and harnesses, demonstrating that action scaling is a promising target for enhancing terminal agent reliability.

By Minki Kang, Ryo Hachiuma, Shaokun Zhang, Subhashree Radhakrishnan, Yonggan Fu, Jindong Jiang, Mingjie Liu, Ehsan Hosseini-Asl, Yi Dong, Yu-Chiang Frank Wang, Byung-Kwan Lee
arXiv Machine Learning
4d ago

Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation

The paper investigates how language‑model judges can make version‑dependent errors when evaluating upgraded agents. Using 35 public coding‑agent submissions, two customer‑service agents, and over a thousand expert‑labeled trajectories, the authors show that fixed judges often reject task‑conditioned error invariance and can incorrectly approve failed patches, especially as agent capability increases. Paired audits of current outputs reduce interval width only marginally, and the study concludes that independent human patch review is still necessary.

By Jiapeng Li
Hugging Face Trending Papers
Jul 30

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts

Reinforcement learning from verifiable rewards (RLVR) for mathematical reasoning suffers from a structural blind spot: on "cliff" prompts-those on which every sampled rollout in a group fails-the group-normalized advantage is identically zero, so GRPO produces no gradient on precisely the prompts at the frontier of the model's capability. We introduce LoRA Scaffolded Policy Optimization (LSPO), a sampling-time mechanism that recovers this lost gradient.

arXiv AI
Sep 2

APEX-EM: Non-Parametric Online Learning for Autonomous Agents via Structured Procedural-Episodic Experience Replay

APEX-EM is a non‑parametric experience memory that stores full procedural‑episodic traces in a typed Procedural Knowledge Graph and retrieves them via semantic search, structural‑signature matching, and graph traversal. It uses a Plan‑Retrieve‑Generate‑Iterate‑Ingest workflow to produce, quality‑gate, and commit experiences, indexing both successes and failures so the agent learns what to reuse and what to avoid. Evaluations on five benchmarks with a shared GPT‑4o backbone show significant performance gains, such as +7.6 pp on BigCodeBench transfer and +1.4 pp on Lifelong Agent Bench, demonstrating that the memory adds to model capability rather than replacing it.

By Pratyay Banerjee, Masud Moshtaghi, Ankit Chadha
arXiv AI
Sep 24

Reinforcement Learning with Decomposed Subtasks

The paper introduces Reinforcement Learning with Decomposed Subtasks (RLDS), a method that splits trajectory rewards into per‑subtask shares before policy updates, replacing the scalar advantage used in Group Relative Policy Optimization (GRPO). RLDS employs Subtask‑Decomposed Advantage Estimation (SDAE) to compute group‑relative advantages and distribute credit to tokens based on subtask importance, focusing on steps where a reflection marks a subtask as consequential. Experiments on four benchmarks—FrozenLake, HotpotQA, ScienceWorld, and DeepResearch—show that RLDS improves performance on high‑heterogeneity tasks (ScienceWorld and FrozenLake) and is more compute‑efficient than scalar GRPO for long rollouts.

By Mattie Terzolo, Mikolaj Sacha, Ayan Sinha, Andrew Rabinovich