Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks
arXiv:2606. 29082v1 Announce Type: cross Abstract: Would experience designing faster GPU kernels also help close in on a long-standing open mathematical conjecture?
arXiv:2608. 09537v1 Announce Type: new Abstract: Foundation world models have made remarkable progress in planning, simulation, and embodied intelligence.
arXiv:2606. 29082v1 Announce Type: cross Abstract: Would experience designing faster GPU kernels also help close in on a long-standing open mathematical conjecture?
The paper introduces Drift-Constrained Optimization (DCO), a framework that treats behavioral drift during fine‑tuning of instruction models as a bounded constraint rather than an uncontrolled side effect. By defining a drift budget, the authors reformulate fine‑tuning as a direction‑selection problem, showing that choosing different update directions can qualitatively change outcomes. Experiments on Qwen3 models demonstrate that carefully selected directions improve scientific reasoning and multilingual translation while preserving reasoning capabilities and general performance.
arXiv:2608. 06714v1 Announce Type: new Abstract: Recent systems for optimizing prompts, programs, and ML workflows typically rely on explicit outer-loop controllers such as evolutionary search, bandits, or textual-gradient methods.
arXiv:2607. 18235v1 Announce Type: cross Abstract: Autonomous discovery systems such as OpenEvolve and TTT-Discover are often used as general-purpose harnesses.
Recent systems for optimizing prompts, programs, and ML workflows typically rely on explicit outer-loop controllers such as evolutionary search, bandits, or textual-gradient methods. We ask a fundamentally different question: how much of this search policy can be internalized by a single tool-using agent?
arXiv:2608. 05651v1 Announce Type: cross Abstract: Large language model (LLM)-driven evolution has shown promise for program search and algorithm discovery, but relying on strong models throughout long evolutionary runs is costly.
arXiv:2608. 09629v1 Announce Type: new Abstract: Self-evolving agents are usually built around prescribed optimization pipelines: the framework decides how to gather evidence, revise a persistent artifact, select candidates, and stop.
arXiv:2607. 06764v1 Announce Type: new Abstract: Recent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (evolutionary search, exhaustive sampling, extended chain-of-thought), or benchmark-specific training in which small models are fine-tuned on ARC data, often with task-specialized architectures.
arXiv:2609.38349v1 Announce Type: cross Abstract: Modern agentic systems combine an AI model with a harness that controls execution and environmental interactions. Harness design strongly affects lon...
arXiv:2606. 02863v1 Announce Type: new Abstract: AI-Driven Research Systems (ADRS) -- systems coupling LLMs with automated evaluation to discover algorithms, proofs, and designs -- are being optimized and adopted across domains, but the tools to analyze them have not kept pace.
The article argues that agentic auto‑research should be guided by dense, intermediate signals of epistemic progress rather than by sparse final benchmarks. It compares this approach to fuzz testing, where coverage provides continuous feedback that directs input mutation. The authors propose controlled experiments to test whether such signals improve discovery efficiency and reduce false positives, and demonstrate in a simulated physics setting that an AI agent using feedback‑driven search uncovers a hidden law while optimization‑driven baselines fail.
The paper introduces Online Surrogate Repair (OSR), a closed‑loop algorithm that decouples the frequency of high‑fidelity evaluations from the length of an agent’s search by selectively updating a surrogate model with sparse, high‑fidelity data. An acquisition rule determines which candidate designs receive expensive evaluations, and the resulting labels refine the surrogate for subsequent episodes. Experiments on synthetic environments and the MADE benchmark show that OSR can reduce regret more efficiently than fixed‑surrogate approaches, requiring fewer oracle queries than high‑fidelity feedback after every episode.