arXiv:2606. 15197v1 Announce Type: cross Abstract: Optimization modeling is inherently hierarchical, requiring a precise sequence of symbolic commitments.
By Jiajun Li, Yu Ding, Shisi Guan, Ran Hou, Wanyuan Wang
MiLoop is a reinforcement‑learning‑based constructive framework for neural combinatorial optimization that propagates selective memory across rollout steps. By fusing current embeddings with historical memory before attention layers and applying adaptive gated updates afterward, it enables a shallow policy to learn dynamic embeddings without external solution labels or search‑space pruning. Experiments on four combinatorial optimization problems show MiLoop consistently generates high‑quality solutions for instances ranging from 100 to 10 million nodes, demonstrating strong generalization.
By Changliang Zhou, Yuanyao Chen, Rongsheng Chen, Zhiyun Lin, Zhenkun Wang
arXiv:2602.13691v2 Announce Type: replace
Abstract: Recent advancements in Large Language Model (LLM) agents have demonstrated strong capabilities in executing complex tasks through tool use. However...
By Yu Li, Guangfeng Cai, Shengtian Yang, Han Luo, Shuo Han, Xu He, Dong Li, Lei Feng
Agentic ESOpt proposes using evolution strategies (ES) instead of reinforcement learning to fine‑tune large language‑model agents for long‑horizon tasks. ES offers model scalability, flexibility, and better long‑horizon credit assignment, enabling full‑parameter optimization with minimal GPU memory. The framework samples parameter perturbations, evaluates agents with rewards, and updates online, achieving notable performance gains on WebArena‑Lite and in test‑time prompt‑parameter co‑evolution.
By Zhi Zheng, Rongsheng Chen, Yunpeng Ba, Zhenkun Wang, Yee Whye Teh, Wee Sun Lee
GRAPE is a two‑stage Bayesian optimization framework that first refines the local gradient posterior using a closed‑form acquisition function and then selects update directions by maximizing expected decrease conditioned on descent. The authors prove that the refinement stage monotonically reduces local uncertainty and that the progress‑aware direction converges to true steepest descent as the posterior sharpens. Empirical results show GRAPE achieves a 5.4× speedup on black‑box adversarial attacks and reduces final average regret by 3.8 log‑units on large language model prompt‑optimization tasks.
By Richard Cornelius Suwandi, Feng Yin
arXiv:2607. 21971v1 Announce Type: new Abstract: Test-time scaling through iterative self-evolution with environment feedback, as demonstrated by AlphaEvolve, shows remarkable performance gains.
By Shujin Wu, Cheng Qian, Xiusi Chen, Heng Ji