arXiv AI

Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents

arXiv:2605. 22148v2 Announce Type: replace Abstract: Self-evolving skill libraries, pioneered by Voyager, let frozen LLM agents accumulate reusable knowledge without weight updates, yet recent evaluation shows that LLM-authored skills deliver $+0.

arXiv AI
Sep 21

Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design

Designer‑RSI presents a continual adaptation framework that lets a frozen frontier model operate professional design software while an external procedural memory learns natural‑language design skills from user traffic. Over five rounds on 1,406 real briefs and 1,869 graded trajectories, the memory grew from 76 to 139 skills, boosting execution success from 72.7% to 99.3% and improving win rates on four design benchmarks. The study shows that widening and deepening the memory, especially together, significantly outperforms a no‑skill baseline.

By Hongyang Du, Lan Yan, Christian Flores, Asim Kadav
arXiv AI
Sep 3

SkillGLoW: Procedural-Family Skill Consolidation for Self-Improving Agents on Long-Horizon Task Streams

SkillGLoW introduces a new way for large language model agents to self‑improve by consolidating procedural skills shared across related tasks. Instead of storing all skills in a single global document or a flat per‑task pool, SkillGLoW aggregates local skills into procedural families, compresses them into de‑instantiated global priors, and regenerates instance‑specific details on demand. Experiments on four diverse benchmarks show that these priors improve performance by an average of 17.2 points over a no‑skill baseline, are more compact than per‑task pools, and enable better transfer to unseen tasks.

By Ao Yan, Xin Zhang, Jiawei Du, Joey Tianyi Zhou
arXiv AI
Aug 25

TRACE: A Self-Evolving Skill Bank for Consistent, Limit-Aware LLM Agents

TRACE (TRAjectory-Contrastive Evolution) is a self‑evolving skill bank that improves the consistency and limit‑awareness of large‑language‑model agents without changing the model weights. By iteratively refining modular skills based on successful and failed trajectories, TRACE raises consistent performance (Pass^3) on the CAR‑bench in‑car assistant tasks from 59.9 % to 94.5 % on GPT‑5.5 and achieves first place on the hidden set with GPT‑5.6‑Sol. The approach demonstrates that a skill‑based, self‑evolution loop can convert a model’s potential into stable, reliable behavior.

By Wenhao Wu, Menghao Zhang, Xin Wang, Zhi Wang, Kun Shao, Jian Luan
arXiv AI
Sep 24

Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents

The paper introduces Just-in-Time Memory (JitMem), a system that defers memory curation until a task is read, allowing a curator to synthesize task‑specific memory payloads based on the current query. Unlike traditional write‑time curation, JitMem retains raw trajectories and trains the curator using immediate task success, avoiding long‑horizon credit‑assignment issues. Experiments on ALFWorld, WebShop, and τ²‑bench show JitMem consistently outperforms both no‑memory agents and existing write‑time memory methods, with improvements of up to 16.3 absolute success‑rate points. whyItMatters":"By curating memory at read time, JitMem enables more effective, task‑adaptive recall that directly improves agent performance across diverse benchmarks."

By Yefan Zhou, Yang Li, Zeyu Leo Liu, Semih Yavuz, Shafiq Joty