arXiv AI By Cheng Li, Jiexiong Liu, Yixuan Chen, Chi Hong

Train What You Deploy:Token-Faithful Post-Training of a Production Coding

Read the original on arXiv AI →

The paper introduces a fidelity‑aware post‑training framework for coding and terminal agents that preserves token and control fidelity by sampling from original prompts, eliminating spurious model calls, and limiting loss computation to verifiable token spans. It also proposes Certified Divergence Proximal Policy Optimization (C‑DPPO), which provides tight two‑sided total variation certification bounds, adaptive‑K rules, budget‑aware sequence guarantees, and error‑robust policy masking. Experiments on Baize5B and Baize10B models show a consistent +3.0‑point performance improvement over standard DPPO, with certificate audits confirming full operational coverage.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 11

T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks

The paper introduces T1, a 122‑billion‑parameter Mixture‑of‑Experts model trained with reinforcement learning to perform long‑horizon terminal tasks such as coding and scientific discovery. T1 operates a real shell in a cloud sandbox, making over 300 tool‑call turns per task and receiving rewards from task‑specific verifiers. The authors detail a training recipe that includes aggressive warm‑starting, TITO construction with drift repair, and rollout‑routing replay, achieving significant performance gains on Terminal‑Bench 2.1 and surpassing GPT‑5.4 and GLM‑5.1 on the Long‑Horizon Terminal Bench.

By Junyao Yang, Yucheng Shi, Zhongzhi Li, Ruhan Wang, Zongxia Li, Haitao Mi, Leowei Liang
arXiv AI
Jun 9

Training-Inference Kernel Contracts: Bounding Divergence in Post-Training and Deployment

arXiv:2606. 07581v1 Announce Type: cross Abstract: A modern post-training pipeline often writes one symbol for its policy, pi_theta, while evaluating it through two different programs: a training kernel optimized for autograd and an inference kernel optimized for low-precision, fused, dynamically batched serving.

By Bruce Changlong Xu, Lan Wu
arXiv Computation and Language
2d ago

CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution

arXiv:2610.10426v1 Announce Type: new Abstract: Terminal-agent capability depends jointly on model weights and the runtime harness that formats prompts, binds tools, and handles error recovery. Exist...

By Jixuan Chen, Jiaxin Zhang, Qinyuan Ye, Yada Pruksachatkun, Haoxiang Zhang, Jingming Zhuo, Yifan Zhang, Yutong Dai, Juntao Tan, Xiangyu Peng, Silvio Savarese, Zeyuan Chen, Lianhui Qin, Chien-Sheng Wu
arXiv AI
Aug 19

LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents

LEGO-RL is a framework that connects native coding-agent harnesses with scalable policy‑gradient training without altering the harnesses’ internal flow. It achieves faithful optimization through in‑process LLM proxying, reliable execution via sandbox orchestration, and observable training with automated validation and a Live UI. Experiments show LEGO‑RL improves the Qwen3.5‑35B‑A3B model’s performance on three native harnesses while preserving high rollout‑training probability correlation.

By Yiming Du, Yuxin Jiang, Tao Yuan, Jianbo Dai, Shaowei Wang, Jierun Chen, Chaofan Tao, Xianzhi Yu, Lifeng Shang, Kam-Fai Wong, Xiaohui Li, Haoli Bai