arXiv Machine Learning

Early-Stopping Thresholds for ES-HyperNEAT: A Data-Driven Approach from Fitness Dynamics

arXiv AI
4d ago

Rank-Reliable Teacher-Guided Fitness Approximation for Expensive Evolutionary Optimization: A TinyML Architecture Search Study

The paper introduces TGL-NSGA-II, a low‑fidelity framework that uses a pretrained teacher to stratify samples by difficulty and class, then applies a short knowledge‑distillation step (KD‑Lite) before scoring candidates on a stratified evaluation set. The teacher‑guided scores are fused with a Gaussian‑process surrogate to select candidates for full evaluation, and the method is evaluated on keyword spotting and bird‑call classification tasks. Results show high Kendall‑τ values (0.74 and 0.62), a 41% reduction in proxy‑score variance, and improved hypervolume and false‑positive rates compared to full NSGA‑II, while running 2.2× faster under a constrained evaluation budget.

By Soumen Garai, Suman Samui
arXiv AI
Sep 2

Bandits in Prod: Hyperparameter Optimization at Inference Time

The paper introduces Online Hyperparameter Optimization (OHPO), framing it as an infinitely many‑armed bandit problem over mixed and conditional search spaces. It proposes the IMABO framework, which couples any bandit policy with any oracle for proposing new configurations, and presents IMOSS—a restart‑free anytime policy with provable regret bounds. Experiments show that IMABO, combined with practical oracles such as TPE, an incumbent‑mutation oracle, and a pretrained tabular foundation model, outperforms random search across a range of settings from classical ML models to LLM‑based agents.

By Louis Abraham, Tuan-Anh Nguyen, Nicolas Devatine
arXiv AI
Sep 2

Flawed in Nature, Perfect through Evolution

arXiv:2609.00129v1 Announce Type: cross Abstract: The performance of artificial intelligence (AI) and machine learning (ML) models degrades when the problem they were trained on drifts. This is a nea...

By J. M. Diederik Kruijssen (Allora Foundation)
arXiv Machine Learning
Jul 1

Predictable GRPO: A Closed-Form Model of Training Dynamics

arXiv:2606. 30789v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has become a standard tool for improving the reasoning ability of large language models, yet its training dynamics are still described empirically: reward trajectories are fit with low-parameter functional forms whose constants carry no mechanistic meaning, and hyperparameter choices remain a matter of trial and error.

By Rajat Ghosh, Datta Nimmaturi, Aryan Singhal, Vaishnavi Bhargava, Henry Wong, Johnu George, Debojyoti Dutta