arXiv Machine Learning

Provably Reduced Sample Cost in Prior-Guided Hyperparameter Optimization

arXiv:2606. 04866v1 Announce Type: new Abstract: Large-scale hyperparameter optimization (HPO) in automated machine learning (AutoML) consumes substantial computational resources, raising growing concerns about scalability and energy efficiency.

arXiv AI
Aug 20

Automated Computational Energy Minimization of ML Algorithms using Constrained Bayesian Optimization

The paper presents a method that uses Constrained Bayesian Optimization (CBO) to minimize the energy consumption of machine learning models while ensuring their generalization performance stays above a specified threshold. By treating energy usage as the primary objective and performance as a constraint, the authors demonstrate that CBO can reduce training energy costs on both regression and classification tasks without sacrificing predictive accuracy.

By Pallavi Mitra, Felix Biessmann
arXiv Machine Learning
Sep 3

HyperMC: Multi-Fidelity Hyperparameter Tuning for Stochastic Gradient MCMC

HyperMC is a multi‑fidelity hyperparameter tuning framework for stochastic gradient Markov chain Monte Carlo (SGMCMC) that combines Hyperband-style resource allocation with kernel Stein discrepancy (KSD) evaluation. It uses successive‑halving brackets to explore a continuous hyperparameter space while progressively refining promising configurations within a fixed computational budget. Robust HyperMC further introduces global grid initialization and elite‑guided local refinement to reduce sensitivity to random candidate generation and noisy evaluations, and theoretical analysis shows that the successive‑halving component selects a near‑optimal configuration with high probability under suitable conditions.

By Ming Tan, Xiyun Jiao
arXiv AI
Sep 2

Efficiently Estimating Optimal Hyperparameter Scaling Laws through Power-Law Entropy Search

The paper introduces Power‑Law Entropy Search (PLES), a computational‑cost‑aware acquisition function that uses multi‑fidelity Bayesian optimization to efficiently estimate optimal hyperparameter scaling laws for large language model training. PLES focuses on reducing the overall uncertainty of scaling law estimates rather than optimizing a single objective, selecting configurations that maximize uncertainty reduction per unit computational cost. Experiments on synthetic benchmarks, surrogate models, and real LLM pre‑training runs show that PLES converges to accurate scaling laws using less than one‑tenth of the computational budget required by conventional grid search and other baselines.

By Zhiliang Chen, Sebastian Ament, David Eriksson, Maximilian Balandat, Eytan Bakshy, Jihao Andreas Lin
Hugging Face Trending Papers
Jul 15

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems

The rapid deployment of machine learning systems across cloud, edge, and enterprise environments has brought model optimization to the forefront of systems-engineering. Despite a rich literature spanning quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference-time optimization, practitioners are often left navigating these techniques through heuristics rather than principled methodology.