arXiv AI By Hanchen Li, Runyuan He, Qizheng Zhang, Changxiu Ji, Qiuyang Mang, Xiaokun Chen, Lakshya A Agrawal, Wei-Liang Liao, Eric Yang, Alvin Cheung, James Zou, Kunle Olukotun, Ion Stoica, Joseph E. Gonzalez

Combee: Scaling Prompt Learning for Self-Improving Language Model Agents

Read the original on arXiv AI →

Combee is a new framework that scales prompt learning for self‑improving language model agents by enabling many agents to run in parallel while learning from their combined traces. It uses parallel scans, an augmented shuffle mechanism, and a dynamic batch size controller to maintain quality and reduce delay. Experiments on AppWorld, Terminal‑Bench, Formula, and FiNER show up to 17× speedup over prior methods with comparable or better accuracy at similar cost.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 30

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization

arXiv:2602. 11351v2 Announce Type: replace Abstract: Proactive large language model (LLM) agents aim to actively plan, query, and interact over multiple turns, enabling efficient task completion beyond passive instruction following and making them essential for real-world, user-centric applications.

By Yihang Yao, Zhepeng Cen, Haohong Lin, Shiqi Liu, Zuxin Liu, Jiacheng Zhu, Zhang-Wei Hong, Laixi Shi, Ding Zhao
arXiv AI
Aug 28

Naive Prompt Optimization: Rethinking the Need for Complex Prompt Search

Naive Prompt Optimization (NPO) is a lightweight, single‑lineage method that iteratively refines prompts using a teacher model’s rollout feedback. It matches or surpasses the performance of more complex optimizers like GEPA while requiring fewer rollouts, and its advantage grows with stronger teacher models. In interactive games, NPO remains competitive, and prompts optimized by NPO transfer well to other student models within the same family.

By Yuan Chang, Xiaoqi Chen