arXiv:2607. 23377v1 Announce Type: cross Abstract: The largest machine learning models in particle physics are also the most expensive to train, yet the return on scaling a given architecture cannot be estimated before that compute is spent.
By Jan-Lucas Uslu, Benjamin Nachman, Christopher Re
The largest machine learning models in particle physics are also the most expensive to train, yet the return on scaling a given architecture cannot be estimated before that compute is spent. Scaling laws have been fit for jets, but none has yet been shown to predict the performance of models it was not fit on.
arXiv:2606. 14870v1 Announce Type: cross Abstract: Foundation models (FMs) trained on large datasets and fine-tuned on downstream tasks have emerged as a powerful paradigm in AI for science.
By Ibrahim Elsharkawy, Joschka Birk, Vinicius Mikuni, Wahid Bhimji, Gregor Kasieczka, Benjamin Nachman
arXiv:2512. 04149v2 Announce Type: replace-cross Abstract: Next token prediction is an attractive pre-training task for jet foundation models, in that it is simulation free and enables excellent generative capabilities that can transfer across datasets.
By Joschka Birk, Anna Hallin, Gregor Kasieczka, Nikol Madzharova, Ian Pang, David Shih
arXiv:2607. 27501v1 Announce Type: new Abstract: We present a lightweight approach to foundation modeling (\textbf{NEXUS}) that leverages pre-trained learning from collider physics data towards out-of-domain tasks in other scientific datasets, using a fully connected autoencoder model with approximately 3 million parameters.
By Liangyu Wu, Qibin Liu, Alexander Yue, Julia Gonski
arXiv:2605. 29283v2 Announce Type: replace-cross Abstract: Recent physics foundation models claim general spatiotemporal forecasting ability, yet their evaluations often collapse performance into a single average score under a fixed training distribution.
By Mengdi Chu, Yang Liu, Ayan Biswas, Han-Wei Shen
The paper introduces a model calibration method using optimal transport to address discrepancies between simulation and experimental data in high-dimensional machine learning applications. Applied to jet tagging in particle physics, the technique calibrates a 128‑dimensional latent representation from a general‑purpose classifier, ensuring downstream derived quantities are properly calibrated. This enables more reliable use of foundation models for jet flavor analysis in LHC experiments and offers a general framework for correcting high‑dimensional simulations across scientific fields.
By Malte Algren, Tobias Golling, Francesco Armando Di Bello, Christopher Pollard
Panda Diplomacy introduces a point‑cloud self‑distillation framework that enables a single foundation‑model architecture and objective to be pre‑trained across three distinct particle‑detector modalities—liquid argon time‑projection chambers, collider TPCs, and water Cherenkov detectors—without extensive modification. Using only 1,000 labeled images for downstream adaptation, the resulting Panda V2 model matches or surpasses specialized baselines that require orders of magnitude more supervision, achieving state‑of‑the‑art particle‑clustering performance with 70× fewer labeled events on sPHENIX and up to 1,000× fewer labels on LArTPC data. Linear probes further demonstrate that the model’s latent space captures physically meaningful structures such as particle causality and track curvature.
By Samuel Young, C\'esar Jes\'us-Valls, Kazuhiro Terao
arXiv:2509. 13805v4 Announce Type: replace-cross Abstract: Foundation models have revolutionized natural language processing through a ``train once, deploy anywhere'' paradigm, where a single pre-trained model adapts to countless downstream tasks without retraining.
By Florian Wiesner, Zo\"e J. Gray, Matthias Wessling, Stephen Baek
arXiv:2606. 19781v1 Announce Type: cross Abstract: Neural scaling laws describe how model performance improves as a power law in compute, model size, and dataset size.
By Jan-Lucas Uslu, Kevin Greif, Daniel Whiteson, Benjamin Nachman
arXiv:2606. 16035v1 Announce Type: cross Abstract: Modern particles physics experiments have demonstrated an increasing need for fast, high-fidelity detector simulation as detector components have improved and subsequent computational requirements approach the limits of available resources.
By Cole Granger, James Giroux, Richard Tyson, Maurizio Ungaro, Cristiano Fanelli
The paper applies sparse autoencoders to a neutrino foundation model trained on IceCube data, uncovering a validated atlas of physical concepts within the model’s internal representation. Causal analysis shows the direction reconstruction head largely ignores this atlas, whereas an uncertainty head trained on the same representation effectively uses quality and brightness features, improving angular resolution from 20.2° to 3.2° at 20% efficiency. These findings demonstrate that mechanistic interpretability can expose latent physics and guide the design of downstream tasks.
By Rapha\"el Bonnet-Guerrini, Johann Ioannou-Nikolaides, Inar Timiryasov, Vincenzo Piuri