Optimize Cheap, Deploy Strong: Cost-Aware Cross-Tier Transfer for Evolutionary Optimization
arXiv:2608. 10694v1 Announce Type: cross Abstract: Evolutionary optimization of LLM prompts and agentic programs (e.
arXiv:2608. 10694v1 Announce Type: cross Abstract: Evolutionary optimization of LLM prompts and agentic programs (e.
arXiv:2606. 10389v1 Announce Type: new Abstract: Recent advances in LLM-driven code evolution have enabled automated discovery by iteratively generating and improving programs.
arXiv:2606. 29119v1 Announce Type: cross Abstract: We introduce a pre-registered screening rule that decides, before any implementation, whether an evolutionary / population / lifecycle outer loop over neural-network parameters or structure is worth building.
arXiv:2609.37808v1 Announce Type: new Abstract: Protein optimization aims to discover high-fitness sequences under a limited experimental budget. Existing machine-learning methods use task-specific p...
arXiv:2606. 03938v1 Announce Type: cross Abstract: Multi-epoch training is becoming the standard now that compute is growing faster than the supply of high-quality text.
arXiv:2609.14418v1 Announce Type: cross Abstract: Dynamic multi-mode resource-constrained project scheduling requires decisions to be made under precedence constraints, limited resources, multiple ex...
The paper introduces TGL-NSGA-II, a low‑fidelity framework that uses a pretrained teacher to stratify samples by difficulty and class, then applies a short knowledge‑distillation step (KD‑Lite) before scoring candidates on a stratified evaluation set. The teacher‑guided scores are fused with a Gaussian‑process surrogate to select candidates for full evaluation, and the method is evaluated on keyword spotting and bird‑call classification tasks. Results show high Kendall‑τ values (0.74 and 0.62), a 41% reduction in proxy‑score variance, and improved hypervolume and false‑positive rates compared to full NSGA‑II, while running 2.2× faster under a constrained evaluation budget.
The paper introduces Online Hyperparameter Optimization (OHPO), framing it as an infinitely many‑armed bandit problem over mixed and conditional search spaces. It proposes the IMABO framework, which couples any bandit policy with any oracle for proposing new configurations, and presents IMOSS—a restart‑free anytime policy with provable regret bounds. Experiments show that IMABO, combined with practical oracles such as TPE, an incumbent‑mutation oracle, and a pretrained tabular foundation model, outperforms random search across a range of settings from classical ML models to LLM‑based agents.
arXiv:2607. 14408v1 Announce Type: new Abstract: A self-evolving agentic loop repeatedly proposes a tweaked version of an agent (its prompt template or program) and accepts or rejects the change based on a per-iteration quality signal.
arXiv:2609.00129v1 Announce Type: cross Abstract: The performance of artificial intelligence (AI) and machine learning (ML) models degrades when the problem they were trained on drifts. This is a nea...
arXiv:2605.29268v3 Announce Type: replace-cross Abstract: LLM-guided evolutionary search (Evolve systems) has reached state-of-the-art results on mathematical and combinatorial tasks, yet most existi...
arXiv:2606. 30789v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has become a standard tool for improving the reasoning ability of large language models, yet its training dynamics are still described empirically: reward trajectories are fit with low-parameter functional forms whose constants carry no mechanistic meaning, and hyperparameter choices remain a matter of trial and error.