Hyperparameter Search with Transformers and Ray Tune
Related stories
Introducing Optimum: The Optimization Toolkit for Transformers at Scale
LiFT: Local Search via Linear Programming for Overfitting-Controlled Transformers
arXiv:2606. 16243v1 Announce Type: new Abstract: This paper proposes a Linear Programming (LP)-based local search framework for fine-tuning pretrained transformer models with explicit control against overfitting.
Accelerated Inference with Optimum and Transformers Pipelines
Retrieval Augmented Generation with Huggingface Transformers and Ray
Fine-Tune ViT for Image Classification with đ¤ Transformers
Train your first Decision Transformer
Statistically Valid Hyperparameter Selection: From Tuning to Guarantees
arXiv:2606. 25601v1 Announce Type: cross Abstract: Hyperparameter selection is a critical step in the deployment of modern artificial intelligence systems, given the need to tune degrees of freedom such as inference-time parameters, implementation-level settings, and thresholds driving decision rules.
Symmetry-Aware Transformer Training for Automated Planning
arXiv:2508. 07743v2 Announce Type: replace Abstract: While transformers excel in many settings, their application in the field of automated planning is limited.
Exploring new directions in enhancing the ACTS parameter optimization suite
The article investigates how Bayesian optimization can improve the ACTS parameter optimization suite for chargedâparticle reconstruction. By comparing Expected Improvement and Upper Confidence Bound with TPE and random search on an eightâparameter problem, extending the best method to fifteen parameters, and applying Expected Hypervolume Improvement for multiâobjective tuning, the study shows that Bayesian acquisition methods find strong configurations earlier and maintain advantages in heldâout validation. The results demonstrate that Bayesian optimization enhances ACTS autoâtuning through more efficient evaluations, broader search spaces, and the ability to select from nonâdominated tradeâoff solutions.
Statistically Valid Post-Training Hyperparameter Selection: From Tuning to Guarantees
The paper introduces a statistical framework for postâtraining hyperparameter selection, emphasizing the learnâthenâtest (LTT) paradigm. It treats hyperparameter tuning as a multiple hypothesis testing problem over a candidate set, enabling the selection of hyperparameters that meet specified reliability constraints such as risk bounds or informationâtheoretic limits. The framework provides finiteâsample control of error probabilities using pâvalues, eâvalues, and concentration inequalities derived from first principles.
The Transformer as a Polar State Estimator
arXiv:2605. 11007v2 Announce Type: replace-cross Abstract: We show that the core components of the Transformer -- attention, residual connections, and normalization -- arise naturally from a single geometric state estimation problem.