The paper investigates whether two particle physics foundation models, OmniLearned and ParticleViT, pretrained on high‑energy proton–proton and electron–proton collisions can transfer knowledge to a low‑energy neutrino experiment. Using MINERvA neutrino–nucleus scattering data, the authors evaluate the models on energy regression and charged‑current pion classification tasks, finding that the pretrained models outperform similarly sized models trained from scratch, with OmniLearned excelling in regression and ParticleViT in classification. When the same transformer architecture is initialized from unrelated text pretraining (BERT), the performance advantage is minimal for classification and absent for regression, indicating that particle‑level foundation models capture inductive biases that generalize across energy scales, detector technologies, and physics processes.
By Gregor Krzmanc, Vinicius Mikuni, Benjamin Nachman, Callum Wilkinson
arXiv:2512. 04149v2 Announce Type: replace-cross Abstract: Next token prediction is an attractive pre-training task for jet foundation models, in that it is simulation free and enables excellent generative capabilities that can transfer across datasets.
By Joschka Birk, Anna Hallin, Gregor Kasieczka, Nikol Madzharova, Ian Pang, David Shih
arXiv:2607. 23377v1 Announce Type: cross Abstract: The largest machine learning models in particle physics are also the most expensive to train, yet the return on scaling a given architecture cannot be estimated before that compute is spent.
By Jan-Lucas Uslu, Benjamin Nachman, Christopher Re
The largest machine learning models in particle physics are also the most expensive to train, yet the return on scaling a given architecture cannot be estimated before that compute is spent. Scaling laws have been fit for jets, but none has yet been shown to predict the performance of models it was not fit on.
arXiv:2606. 19781v1 Announce Type: cross Abstract: Neural scaling laws describe how model performance improves as a power law in compute, model size, and dataset size.
By Jan-Lucas Uslu, Kevin Greif, Daniel Whiteson, Benjamin Nachman
The paper introduces a model calibration method using optimal transport to address discrepancies between simulation and experimental data in high-dimensional machine learning applications. Applied to jet tagging in particle physics, the technique calibrates a 128‑dimensional latent representation from a general‑purpose classifier, ensuring downstream derived quantities are properly calibrated. This enables more reliable use of foundation models for jet flavor analysis in LHC experiments and offers a general framework for correcting high‑dimensional simulations across scientific fields.
By Malte Algren, Tobias Golling, Francesco Armando Di Bello, Christopher Pollard