arXiv AI

WMLLM: Self-Evolving Optimization Agents via Predict-Then-Act World Modeling

WMLLM is a self‑evolving optimization‑agent framework that uses a predict‑then‑act world modeling approach. The agent first predicts promising optimization directions and then generates candidates, refining both its world model and strategy through multi‑turn refinement, population‑based search, and reinforcement learning. Experiments on black‑box tasks, particularly multi‑objective molecular optimization, demonstrate that WMLLM improves sample efficiency and achieves state‑of‑the‑art results within a limited evaluation budget.

arXiv AI
Jun 6

MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery

arXiv:2606. 06473v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly applied to long-horizon tasks such as scientific discovery and machine learning engineering (MLE), where sustained self-evolution becomes a key capability.

By Shangheng Du, Xiangchao Yan, Jinxin Shi, Zongsheng Cao, Shiyang Feng, Zichen Liang, Boyuan Sun, Tianshuo Peng, Yifan Zhou, Xin Li, Jie Zhou, Liang He, Bo Zhang, Lei Bai
arXiv AI
Aug 28

MCCE: A Framework for Multi-LLM Collaborative Search in Discrete Spaces with Similarity-Filtered Preference Learning

The paper introduces MCCE, a hybrid framework that combines a frozen closed‑source large language model (LLM) with a lightweight, trainable model for multi‑objective discrete optimization. By maintaining a trajectory memory and refining the small model through reinforcement learning, the two models jointly enhance global exploration and learning. Experiments on drug‑design benchmarks demonstrate that MCCE achieves state‑of‑the‑art Pareto front quality, outperforming existing baselines.

By Nian Ran, Zhongzheng Li, Yue Wang, Qingsong Ran, Xiaoyuan Zhang, Shikun Feng, Richard Allmendinger, Xiaoguang Zhao
arXiv Machine Learning
Aug 19

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

Agentic ESOpt proposes using evolution strategies (ES) instead of reinforcement learning to fine‑tune large language‑model agents for long‑horizon tasks. ES offers model scalability, flexibility, and better long‑horizon credit assignment, enabling full‑parameter optimization with minimal GPU memory. The framework samples parameter perturbations, evaluates agents with rewards, and updates online, achieving notable performance gains on WebArena‑Lite and in test‑time prompt‑parameter co‑evolution.

By Zhi Zheng, Rongsheng Chen, Yunpeng Ba, Zhenkun Wang, Yee Whye Teh, Wee Sun Lee
arXiv Machine Learning
Sep 7

Small Molecule Optimization with Large Language Models

The paper introduces Mol-E, an evolutionary algorithm that leverages large language models trained on molecular data to generate candidate molecules. Mol-E achieves state‑of‑the‑art performance on the Practical Molecular Optimization benchmark, scoring 17.500 in the task‑agnostic regime and 20.551 in the task‑informed regime. It also outperforms baseline methods in multi‑property optimization tasks involving docking against DRD2, MK2, and AChE.

By Philipp Guevorguian, Menua Bedrosian, Tigran Fahradyan, Gayane Chilingaryan, Armen Aghajanyan, Hrant Khachatrian
arXiv AI
Sep 4

CoMAP: Co-Evolving World Models and Agent Policies for LLM Agents

CoMAP introduces a framework that jointly evolves textual world models and agent policies through a closed‑loop interaction. At each decision step the world model forecasts future state feedback for candidate actions, while the agent reflects on the reliability of this feedback to refine its action. The resulting on‑policy trajectories are used to self‑distill and update the world model, improving prediction accuracy and long‑horizon decision‑making across embodied planning, web navigation, and tool‑use benchmarks.

By Youwei Liu, Jian Wang, Hanlin Wang, Wenjie Li