Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
Read the original on arXiv AI →The paper introduces SORL, a framework that stabilizes off‑policy reinforcement learning for long‑horizon large language model agents. It identifies two key instability sources—token‑level policy granularity mismatches and high‑variance off‑policy updates—and proposes turn‑level importance sampling and clipping‑triggered normalization to align optimization with multi‑turn interactions. Two instantiations, SO‑PPO and SO‑GRPO, are evaluated on open‑domain, multi‑hop, and medical QA benchmarks, as well as on asynchronous RL for mathematical reasoning, showing improved robustness without the need for early stopping or heuristic tuning.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.