The paper introduces a statistical framework for post‑training hyperparameter selection, emphasizing the learn‑then‑test (LTT) paradigm. It treats hyperparameter tuning as a multiple hypothesis testing problem over a candidate set, enabling the selection of hyperparameters that meet specified reliability constraints such as risk bounds or information‑theoretic limits. The framework provides finite‑sample control of error probabilities using p‑values, e‑values, and concentration inequalities derived from first principles.
By Amirmohammad Farzaneh, Osvaldo Simeone
arXiv:2211. 14411v5 Announce Type: replace-cross Abstract: Hyperparameter optimization (HPO) is crucial for strong performance of deep learning algorithms and real-world applications often impose some constraints, such as on memory usage or latency, on top of the performance requirement.
By Shuhei Watanabe, Frank Hutter
arXiv:2604. 13130v2 Announce Type: replace Abstract: We study learning to learn through the lens of hyperparameter tuning.
By Saumya Goyal, Rohith Rongali, Ritabrata Ray, Barnab\'as P\'oczos
arXiv:2606. 15569v1 Announce Type: new Abstract: Test-time training (TTT) adapts a pretrained model to each prompt via parameter updates, improving accuracy under pretraining-to-test distribution shifts.
By Tomoya Wakayama
HyperMC is a multi‑fidelity hyperparameter tuning framework for stochastic gradient Markov chain Monte Carlo (SGMCMC) that combines Hyperband-style resource allocation with kernel Stein discrepancy (KSD) evaluation. It uses successive‑halving brackets to explore a continuous hyperparameter space while progressively refining promising configurations within a fixed computational budget. Robust HyperMC further introduces global grid initialization and elite‑guided local refinement to reduce sensitivity to random candidate generation and noisy evaluations, and theoretical analysis shows that the successive‑halving component selects a near‑optimal configuration with high probability under suitable conditions.
By Ming Tan, Xiyun Jiao
arXiv:2602. 05786v3 Announce Type: replace Abstract: Tree-boosting is a widely used machine learning technique for tabular data.
By Floris Jan Koster, Fabio Sigrist
arXiv:2401.03580v2 Announce Type: replace-cross
Abstract: We study hyperparameter optimization (HPO) from a numerical-optimization perspective and propose a multi-objective, damped Gauss--Newton sear...
By Qinwu Xu
The paper explores when different decision proxies are appropriate for decision‑focused learning (DFL) in optimization problems with uncertainty. It identifies problem properties that justify using a particular proxy and proposes alternative proxies that maintain learning complexity. Experiments on continuous, discrete, and objective‑ or constraint‑uncertain problems demonstrate the effectiveness of these approaches.
By Noah Schutte, Grigorii Veviurko, Krzysztof Postek, Neil Yorke-Smith
arXiv:2608.20998v1 Announce Type: new
Abstract: Reservoir computing (RC) couples a fixed recurrent dynamical system with a trained lightweight readout, but this efficiency is partly lost during hyper...
By Sara Malacarne, Andrea Ceni, Claudio Gallicchio
arXiv:2608. 13793v1 Announce Type: cross Abstract: Machine learning (ML) has become an indispensable part of modern engineering design workflows.
By Tyler R. Johnson, Kian Ben-Jacob, Christopher P. Muller, Ramin Bostanabad
arXiv:2412. 12807v4 Announce Type: replace-cross Abstract: Selective classification is a powerful tool for automated decision-making in high-risk scenarios, allowing classifiers to act only when confident and abstain when uncertainty is high.
By Mohamed Ndaoud, Peter Radchenko, Bradley Rava
The paper proposes a Bayesian decision framework for multiobjective optimization under uncertainty, focusing on maximizing the expected hypervolume over a finite set of input points. It demonstrates that gradient‑based stochastic optimization can be applied, especially when dominated points are handled carefully, and suggests using Gaussian Processes as differentiable surrogate models when direct gradients are unavailable. Additionally, the authors introduce active learning strategies via acquisition functions to build surrogate models tailored to the multiobjective problem and evaluate these strategies on simple analytical benchmarks.
By Victor Trappler (Mines Saint-\'Etienne MSE, LIMOS, FAYOL-ENSMSE, FAYOL-ENSMSE)