SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution
arXiv:2607. 26784v1 Announce Type: new Abstract: Large language model agents often encounter related yet distinct tasks that share reusable solution patterns.
SkillRL is a framework that enhances large language model agents by automatically discovering and evolving skills from raw experience. It builds a hierarchical skill library called SkillBank, uses an adaptive retrieval strategy for heuristics, and allows the skill library to co‑evolve with the agent’s policy during reinforcement learning. These techniques reduce token usage and improve reasoning, achieving state‑of‑the‑art results on ALFWorld, WebShop, and seven search‑augmented tasks, outperforming baselines by 15.3% and remaining robust as task complexity grows.
arXiv:2607. 26784v1 Announce Type: new Abstract: Large language model agents often encounter related yet distinct tasks that share reusable solution patterns.
SkillGraph introduces a skill library framework that models reusable skills as nodes in a directed graph, with typed edges representing prerequisite, enhancement, and co-occurrence relationships. When presented with a new task, the system retrieves an ordered subgraph of relevant skills, guiding multi-step decision making. The graph is continuously refined through agent trajectories and reinforcement learning, enabling simultaneous improvement of the skill library and the agent policy, and achieving state‑of‑the‑art results on ALFWorld, WebShop, and several search‑augmented QA tasks.
arXiv:2606. 03692v1 Announce Type: new Abstract: Recent AI agents can flexibly invoke skills to solve complex tasks, but their long-term improvement is fundamentally constrained by a lack of systematic skill construction, accumulation, and transfer.
Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet standard agentic reinforcement learning treats tasks as independent episodes, while existing approaches to skill learning either focus on repeated attempts of one task or use pipelines with multiple stages that entangle extraction, retrieval, and execution.
APEx is a hierarchical framework that organizes a deep research agent’s interaction history into instance-level trajectory memories and category-level procedural skills. It couples these through an Executor, Distiller, and Planner, trained with a three-stage alternating GRPO paradigm to enable reward-guided skill distillation. At test time, distilled skills act as procedural priors for online Planner adaptation via skill-guided reinforcement learning, achieving state‑of‑the‑art results on seven benchmarks, outperforming GPT‑5.4 by 14.7 points and the best memory‑augmented baseline by 3.0 points.
arXiv:2608. 15071v1 Announce Type: new Abstract: Learning from experience is critical for developing capable, self-improving large language model (LLM) agents.
arXiv:2608. 15165v1 Announce Type: new Abstract: Large language model (LLM) agents can continually improve without parameter updates by converting historical experience into reusable procedural knowledge.
arXiv:2606. 07603v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit strong reasoning capabilities, yet most LLM-based agents are statically deployed and unable to improve through task interactions.
arXiv:2608.30760v1 Announce Type: new Abstract: Recent studies have shown that multimodal large language models (MLLMs) can serve as embodied agents, translating language instructions and visual obse...
CODESKILL is an LLM-based framework that learns to extract, evolve, and maintain procedural skills from coding-agent trajectories. It treats skill extraction and skill-bank management as a learnable policy trained with reinforcement learning, using a hybrid reward combining rubric-based skill quality and verifiable execution feedback. Experiments on EnvBench, SWE-Bench Verified, and Terminal-Bench 2 demonstrate that CODESKILL raises average pass rates by 11.03 over a no-skill baseline and by 5.10 over the strongest prompt-based or memory baseline while keeping a compact skill bank.
arXiv:2607. 13854v2 Announce Type: replace Abstract: Multimodal agents that think with images iteratively manipulate visual evidence and invoke tools across many steps.
arXiv:2605. 28390v2 Announce Type: replace Abstract: Test-time skill evolving is regarded as a new paradigm for enhancing deployed agentic systems.