arXiv AI

Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries

arXiv:2605. 19576v2 Announce Type: replace Abstract: Self-evolving skill libraries face a silent failure mode we term \emph{library drift}: unbounded skill accumulation without outcome-driven lifecycle management causes retrieval degradation, false-positive injections, and performance stagnation.

arXiv AI
Aug 7

When Self-Evolution Backfires: Pre-Commit Gating against Skill Contamination in LLM Agents

arXiv:2608. 05810v1 Announce Type: new Abstract: Self-evolving agents accumulate capability by distilling reusable skills from their execution trajectories, but we find this process is not monotonic: past a critical pool size, newly added skills degrade performance instead of improving it.

By Linfang Shang, Ming Xu, Yiding Sun, Tianle Xia, Lingxiang Hu, Lan Xu, Ning Zheng
arXiv AI
Sep 7

From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use Agents

The paper introduces an online skill‑evolution framework that transforms interaction traces and evaluator feedback into a persistent, versioned library of reusable procedures for computer‑use agents. By executing each iteration against a frozen library snapshot, the system updates skills without altering the underlying model parameters. Experiments across four OSWorld domains show that the evolving library consistently outperforms an empty‑library baseline, with gains ranging from 5.7 to 18.6 percentage points, while also revealing domain‑specific temporal stability and challenges in skill retrieval and revision.

By Longtao Hu, Xiao Liang, Linchao Zhu
arXiv Computation and Language
2d ago

TRACE: Diagnosing Verifier Brittleness in Agentic Evaluation

The paper introduces TRACE, a protocol designed to diagnose whether changes in verifier scores for large language model agents reflect actual changes in agent behavior or merely alterations in the evaluation process. TRACE works by applying targeted changes to evaluation components, running paired experiments, and re‑scoring unchanged trajectories to isolate the source of score variation. Experiments on synthetic tasks and public benchmark tasks demonstrate that seemingly significant score shifts can often be attributed to evaluation artifacts, while TRACE can also detect genuine behavioral changes.

By Radhika Gaonkar