arXiv:2606. 03108v1 Announce Type: new Abstract: Autonomous LLM training is often framed as recipe search, which leaves the training harness largely static.
By Guhong Chen, Yingcheng Shi, Yongbin Li, Binhua Li, Xander Xu, Hu Wei, Shiwen Ni, Min Yang, Jieping Ye
COBRA‑Skills is a new framework that treats skill optimization for large language model agents as a budgeted sequential problem over a dynamically evolving candidate set. It uses contextual‑bandit prioritization to focus evaluations on promising or informative candidates and refines the skill population based on execution feedback. In experiments across six agent benchmarks and three target models, COBRA‑Skills outperforms existing methods, cuts optimization cost by 55–58 % compared to SkillOpt, and requires only 50 unique optimization examples per benchmark.
By Pingchen Lu, Xiangyi Wang, Xiang Li, Jie Mao, Zikun Qu, Junfeng Luo, Yao Shu, Bryan Kian Hsiang Low, Zhongxiang Dai
arXiv:2608. 05628v1 Announce Type: new Abstract: Although agent skills equip LLMs with reusable procedural knowledge, manual maintenance suffers from high costs, unscalability, and misalignment.
By Yuru Feng, Yaoqi Chen, Beidi Zhao, Qianxi Zhang, Xinjiang Wang, Jianan Lu, Zhirui Wang, Shusen Xu, Zengzhong Li, Qi Chen
arXiv:2607. 22688v1 Announce Type: new Abstract: Post-training agents for automated AI research requires optimizing not only model parameters, but also the runtime harness that shapes how research trajectories are generated, evaluated, and learned from.
By Zhengyu Chen, Teng Xiao, Huaisheng Zhu, Yige Yuan, Luan Zhang, Jingang Wang
Although agent skills equip LLMs with reusable procedural knowledge, manual maintenance suffers from high costs, unscalability, and misalignment. Real-world deployments thus require autonomous, on-demand skill evolution at test time, constrained by limited interaction budgets and a lack of training or validation sets.
The paper introduces CHART, a curriculum that rotates harnesses during training to teach search agents parallel search strategies robustly across different harness configurations. Unlike static harness augmentation, CHART gradually consolidates behavior by graduating learned harnesses and replacing them, maintaining a reward gap that drives learning. Experiments show CHART enables agents to parallelize on 89% of held‑out harnesses, improves performance on a new QA task by 5.6pp, and benefits more from meta‑harness search than baselines.
By Xinlu Zhang, Ying-Chun Lin, Zhihan Zhang, Besnik Fetahu, Xi Chen