arXiv AI By Yu Gu, Zhi Zheng, Yunpeng Ba, Xialiang Tong, Mingxuan Yuan, Zhenkun Wang

Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging

Read the original on arXiv AI →

arXiv:2608. 05541v1 Announce Type: new Abstract: Evolution Strategy (ES) is a promising alternative to gradient-based fine-tuning for resource-constrained Large Language Model (LLM) reasoning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 28

Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO

The paper investigates Evolution Strategies (ES) as a memory‑efficient post‑training method for large language model (LLM) reasoning. It demonstrates that ES outperforms Group Relative Policy Optimization (GRPO) by achieving broader reasoning coverage, improving Pass@K metrics, and avoiding entropy collapse. The study also reveals that ES’s performance gains stem from sparse, high‑magnitude parameter updates, do not cause catastrophic forgetting, and can be combined with GRPO in a sequential training strategy.

By Yunpeng Ba, Zhi Zheng, Yue Xie, Jiaqing Li, Xialiang Tong, Tao Zhong, Mingxuan Yuan, Zhichao Lu, Xuyang Wu, Zhenkun Wang
arXiv AI
Sep 15

Drift-Constrained Optimization: Only Direction Matters in Fine-Tuning Instruct Models

The paper introduces Drift-Constrained Optimization (DCO), a framework that treats behavioral drift during fine‑tuning of instruction models as a bounded constraint rather than an uncontrolled side effect. By defining a drift budget, the authors reformulate fine‑tuning as a direction‑selection problem, showing that choosing different update directions can qualitatively change outcomes. Experiments on Qwen3 models demonstrate that carefully selected directions improve scientific reasoning and multilingual translation while preserving reasoning capabilities and general performance.

By Fei Yuan, Changjiang Gao, Yilei Tu, Yifeng Liu, Shujian Huang, Yu Qiao