Survival is the Only Reward: Sustainable Self-Training Through Environment-Mediated Selection
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2509. 03206v2 Announce Type: replace-cross Abstract: Learning from reward functions and imitation learning of demonstrations are the two principal approaches for training autonomous systems that interact with an environment through action and observation.
The paper introduces a co‑evolving framework where a target agent improves by learning from its own failures, and a separate failure agent is trained to generate hard negative trajectories. These hard negatives, derived from plausible but incorrect attempts, help the target agent better distinguish successful behavior from subtle errors. Experiments on online shopping, scientific reasoning, and interactive SQL querying show a 5.7% average reward improvement over baseline methods.
arXiv:2609.35897v1 Announce Type: new Abstract: The pursuit of recursive self-improvement (RSI) toward general intelligence is divided between macro-level language model scaling and the interaction-d...
arXiv:2609.17325v1 Announce Type: new Abstract: Biological cells can be viewed as individual, interacting agents whose collective dynamics give rise to adaptive behaviour at multiple levels of organi...
arXiv:2607. 21971v1 Announce Type: new Abstract: Test-time scaling through iterative self-evolution with environment feedback, as demonstrated by AlphaEvolve, shows remarkable performance gains.
SPADE (Self-Play in Adaptive Synthetic Executable Environments) is a reinforcement‑learning framework where a single large language model acts as both an Environment Designer—creating executable, long‑horizon training environments—and a Reasoning Agent—learning to act within those environments. The framework uses a regret signal based on the difference between rewarded performance with and without privileged hints to guide the Designer toward environments that are challenging yet solvable. Experiments show that, when scaled to 30‑billion‑parameter models, SPADE outperforms fixed‑environment baselines by significant margins across math, science, code, and reasoning benchmarks, and improves tool‑use performance on BFCL‑v4 and ACEBench‑Agent. whyItMatters":"By making environment design a learnable component, SPADE enables continuous self‑improvement and demonstrates that adaptive, self‑generated training environments can substantially boost language‑model performance across diverse tasks."