EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning
arXiv:2606. 17680v1 Announce Type: new Abstract: Reinforcement learning (RL) has emerged as a powerful paradigm for training Large Language Models (LLMs) as agents.
arXiv:2606. 17680v1 Announce Type: new Abstract: Reinforcement learning (RL) has emerged as a powerful paradigm for training Large Language Models (LLMs) as agents.
arXiv:2609.01597v1 Announce Type: cross Abstract: Natural language is emerging as a primary feedback channel for improving language agents, capable of conveying intent, preferences, and causal struct...
arXiv:2607.08837v4 Announce Type: replace-cross Abstract: Exploration is essential to RL since a policy cannot improve by repeatedly sampling the behaviors it already prefers. Standard methods inject...
The paper explores Retrospection-Only Fine-Tuning (ROFT), a method where a language-model agent improves its behavior by generating and training on explanations of its own experiences, without external teachers or reward signals. In software‑engineering tasks with Qwen3.5‑4B, ROFT achieves comparable or better solve rates than GRPO while requiring fewer updates and training time, and can learn from failures alone. Behavioral analysis shows ROFT indirectly assigns credit to actions and can produce shorter, more direct solutions when prompted to focus on direct solutions.
arXiv:2607. 09042v1 Announce Type: new Abstract: Reinforcement learning (RL) is increasingly used to post-train vision-language-action (VLA) models, but every update consumes robot rollouts that are slow and costly to collect, making sample efficiency a central concern.
arXiv:2608. 04788v1 Announce Type: cross Abstract: Large language model agents are commonly trained through reinforcement learning with sparse trajectory-level rewards, which offer limited guidance on how strongly individual tokens should be updated.
PlanPO introduces a group planning-aware policy optimization method for multi-turn agentic large language models, addressing the issue of advantage collapse caused by treating all successful trajectories equally. By incorporating coarse-to-fine advantage signals that reflect differences in trajectory and turn lengths, PlanPO encourages agents to learn generalizable planning and generation behaviors. Experiments show a 27.2% average improvement over GRPO on benchmarks such as ALFWorld, WebShop, and SciWorld, with minimal extra training cost.
arXiv:2510.15047v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) as agents often fail to improve in new environments. We identify and characterize a failure mode we call explora...
arXiv:2607. 01754v1 Announce Type: new Abstract: On-policy exploration is a crucial component for training robust Vision-Language Navigation agents, as it exposes the policy to a broader state distribution.
arXiv:2607. 08837v1 Announce Type: cross Abstract: Exploration is essential to RL since a policy cannot improve by repeatedly sampling the behaviors it already prefers.
arXiv:2606. 27136v1 Announce Type: new Abstract: For LLM agents in multi-step interactive environments, a key challenge is to make effective use of accumulated interaction experience.
arXiv:2609.24289v1 Announce Type: new Abstract: As Large Language Model (LLM) agents are applied in continuously interactive environments, driving the evolution of their own capabilities becomes a co...