From Search to Signal: Online Post-Training in Automatic Heuristic Design
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2609. 37968v1 Announce Type: new Abstract: Advances in the coding capabilities of LLM agents allow them to inspect and modify their own instructions, tools, and execution procedures.
arXiv:2607. 13389v1 Announce Type: new Abstract: Reinforcement Learning (RL) post-training is increasingly used to adapt foundation models for reasoning, planning, and feedback-driven robot-learning pipelines, but constrained post-training resources are often summarized by a single total FLOP budget.
arXiv:2608. 16156v1 Announce Type: new Abstract: Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, making fine-grained credit assignment across multi-step interactions difficult.
The study investigates whether large language models (LLMs) can automatically improve the harnesses used for hardware design verification. By evolving harnesses around a fixed subject model on 12 proprietary root‑cause localization tasks, the researchers found that automatically evolved harnesses increased completed attempts by 71‑76% and task coverage by 80‑100%, though overall correct attempts improved only 18‑24%. Despite these gains, the evolved harnesses did not consistently consolidate into a single dominant solution across tasks and metrics, and a separate cross‑benchmark case showed that a repair harness could yield a 35.6% increase in functional passes over a baseline.
arXiv:2603. 23420v2 Announce Type: replace Abstract: If autoresearch is itself a form of research, then autoresearch can be applied to research itself.
arXiv:2606. 20002v1 Announce Type: cross Abstract: This work presents a general framework for training large language models (LLMs) to "Connect the Dots" (CoD), a meta-capability required by long-lifecycle agents: as an LLM-based AI agent gets deployed in an environment, it solves a long sequence of tasks while continuously exploring the environment, learning from its own experiences, and iteratively self-updating its context about the environment, thereby achieving progressively better performance on future tasks conditioned on the updated context.