Hugging Face Trending Papers

Relay, Don't Route: Adaptive Population Handoff for Cost-Efficient LLM-Driven Evolution

Read the original on Hugging Face Trending Papers →

Large language model (LLM)-driven evolution has shown promise for program search and algorithm discovery, but relying on strong models throughout long evolutionary runs is costly. A natural alternative is to combine cheap and strong models under a fixed inference budget.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Sep 12

COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization

COBRA‑Skills is a new framework that treats skill optimization for large language model agents as a budgeted sequential problem over a dynamically evolving candidate set. It uses contextual‑bandit prioritization to focus evaluations on promising or informative candidates and refines the skill population based on execution feedback. In experiments across six agent benchmarks and three target models, COBRA‑Skills outperforms existing methods, cuts optimization cost by 55–58 % compared to SkillOpt, and requires only 50 unique optimization examples per benchmark.

By Pingchen Lu, Xiangyi Wang, Xiang Li, Jie Mao, Zikun Qu, Junfeng Luo, Yao Shu, Bryan Kian Hsiang Low, Zhongxiang Dai