arXiv AI

TeleTune: Evolving Agent Skills From Offline Telemetry

arXiv AI
Sep 28

SkillFlow: Scalable and Efficient Agent Skill Retrieval System

SkillFlow is an open, multi-stage retrieval system that helps AI agents selectively load relevant skills from a large library of community-contributed SKILL.md definitions. The pipeline uses dense retrieval, two rounds of cross-encoder reranking, and LLM-based selection to balance recall and precision. Evaluations on SkillsBench and Terminal-Bench show that SkillFlow improves performance when high-quality skills are available, but retrieval alone does not help if the corpus lacks executable skills for the target domain.

By Fangzhou Li, Pagkratios Tagkopoulos, Ilias Tagkopoulos
arXiv AI
Sep 30

SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation

SkillGym is an automatic pipeline that generates verifiable environments for training skill-use agents. It crawls internet skills, filters for reproducible workflows, and uses a builder‑reviewer process to create difficulty‑controlled tasks with reference solutions and verifiers. The system builds 6.8k environments, collects 19k successful trajectories, and fine‑tunes LLMs from 2B to 122B parameters, improving performance and skill invocation rates.

By Renxi Wang, Mingshan Hee, Fajri Koto, Timothy Baldwin, Haonan Li
arXiv AI
3d ago

AMBER: Training Long-Horizon Web Agents through Append-Only Memory

The paper introduces AMBER, an append‑only memory framework for language‑model agents that interact over long horizons. AMBER lets agents jointly learn to reason, act, and write free‑form memory, guaranteeing retention by construction and enabling end‑to‑end reinforcement learning without extensive curated data. Experiments on WebArena Lite show AMBER outperforms overwrite‑based memory by 4.09 percentage points in average success and improves task completion rates in repeated runs.

By Chinmay Savadikar, Zhaoyu Zhang, Mingyu Zhao, Shuang Xie, Han Li, Tianfu Wu, Lingyun Wang
arXiv AI
Aug 3

Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember

arXiv:2607. 29468v1 Announce Type: new Abstract: Self-play agents can generate training problems without questions from target benchmarks, but their curricula lack persistent state: failures affect gradients yet do not explicitly shape future practice.

By Zenghuang Fu, Zhaoyang Li, Qiuyuan Ai, Haoyu Wu, Minghui Wu, Chenxu Zhao, Ante Wang, Guannan He, Changwei Wang