arXiv AI By Zhongzheng Li, Qingsong Ran, Shikun Feng, Nian Ran, Wenhao Li, Xiaoyuan Zhang, Yue Wang, Xiaoguang Zhao

WMLLM: Self-Evolving Optimization Agents via Predict-Then-Act World Modeling

Read the original on arXiv AI →

WMLLM is a self‑evolving optimization‑agent framework that uses a predict‑then‑act world modeling approach. The agent first predicts promising optimization directions and then generates candidates, refining both its world model and strategy through multi‑turn refinement, population‑based search, and reinforcement learning. Experiments on black‑box tasks, particularly multi‑objective molecular optimization, demonstrate that WMLLM improves sample efficiency and achieves state‑of‑the‑art results within a limited evaluation budget.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 6

MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery

arXiv:2606. 06473v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly applied to long-horizon tasks such as scientific discovery and machine learning engineering (MLE), where sustained self-evolution becomes a key capability.

By Shangheng Du, Xiangchao Yan, Jinxin Shi, Zongsheng Cao, Shiyang Feng, Zichen Liang, Boyuan Sun, Tianshuo Peng, Yifan Zhou, Xin Li, Jie Zhou, Liang He, Bo Zhang, Lei Bai
arXiv AI
Aug 28

MCCE: A Framework for Multi-LLM Collaborative Search in Discrete Spaces with Similarity-Filtered Preference Learning

The paper introduces MCCE, a hybrid framework that combines a frozen closed‑source large language model (LLM) with a lightweight, trainable model for multi‑objective discrete optimization. By maintaining a trajectory memory and refining the small model through reinforcement learning, the two models jointly enhance global exploration and learning. Experiments on drug‑design benchmarks demonstrate that MCCE achieves state‑of‑the‑art Pareto front quality, outperforming existing baselines.

By Nian Ran, Zhongzheng Li, Yue Wang, Qingsong Ran, Xiaoyuan Zhang, Shikun Feng, Richard Allmendinger, Xiaoguang Zhao
arXiv Machine Learning
Aug 19

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

Agentic ESOpt proposes using evolution strategies (ES) instead of reinforcement learning to fine‑tune large language‑model agents for long‑horizon tasks. ES offers model scalability, flexibility, and better long‑horizon credit assignment, enabling full‑parameter optimization with minimal GPU memory. The framework samples parameter perturbations, evaluates agents with rewards, and updates online, achieving notable performance gains on WebArena‑Lite and in test‑time prompt‑parameter co‑evolution.

By Zhi Zheng, Rongsheng Chen, Yunpeng Ba, Zhenkun Wang, Yee Whye Teh, Wee Sun Lee