arXiv:2606. 14373v1 Announce Type: cross Abstract: The workflow from particle collision to physics analysis passes through a series of reconstruction steps that are traditionally modular and disconnected, with no shared representation linking low-level detector data to high-level analysis tasks.
By Farouk Mokhtar, Joosep Pata, Michael Kagan, Javier Duarte
arXiv:2412. 10665v3 Announce Type: replace-cross Abstract: We introduce a foundation model for event classification in high-energy physics, built on a Graph Neural Network architecture and trained on 120 million simulated proton-proton collision events spanning 12 distinct physics processes.
By Joshua Ho, Benjamin Ryan Roberts, Shuo Han, Haichen Wang
arXiv:2606. 14813v1 Announce Type: cross Abstract: Jet tagging at the Large Hadron Collider increasingly relies on deep learning models trained on massive simulated datasets, leading to high computational costs and limited robustness to detector mismodeling.
By Guillaume Letellier (LPCC), Antonin Vacheret (LPCC), Fr\'ed\'eric Jurie
arXiv:2608. 14278v1 Announce Type: cross Abstract: We present Pairton, an iterative framework for reconstructing short-lived particles in high-energy collision events.
By Andreas Hermansen, Chris Scheulen, Tobias Golling
arXiv:2606. 20437v1 Announce Type: cross Abstract: Charged-particle tracking -- reconstructing trajectories from sparse detector measurements -- is a fundamental high-energy-physics inference problem and a canonical example of learning under extreme combinatorial ambiguity.
By Siqi Miao, Shitij Govil, Jack P. Rodgers, Mia Liu, Javier Duarte, Shih-Chieh Hsu, Yuan-Tang Chou, Pan Li
The paper introduces a Deep Sets surrogate for optimal transport (OT) that respects key metric properties—non-negativity, exchange symmetry, and zero self-distance—while leaving the triangle inequality unconstrained. Applied to the Energy Mover's Distance between collider events, the Metric-Aware Particle Flow Network achieves percent‑level mean absolute percentage error and markedly higher inference throughput compared to other exact and approximate methods. The architectural constraints also dramatically reduce triangle‑inequality violations, improving geometric fidelity across a large set of held‑out event triplets.
By Lauren Hay, Rishabh Jain, Matt LeBlanc, Jennifer Roloff
arXiv:2609.17738v1 Announce Type: cross
Abstract: Many self-supervised methods for training foundation models at the Large Hadron Collider (LHC) rely on data augmentations to encourage the model to e...
By Ho Fung Tsoi, Dylan Rankin
The paper introduces a model calibration method using optimal transport to address discrepancies between simulation and experimental data in high-dimensional machine learning applications. Applied to jet tagging in particle physics, the technique calibrates a 128‑dimensional latent representation from a general‑purpose classifier, ensuring downstream derived quantities are properly calibrated. This enables more reliable use of foundation models for jet flavor analysis in LHC experiments and offers a general framework for correcting high‑dimensional simulations across scientific fields.
By Malte Algren, Tobias Golling, Francesco Armando Di Bello, Christopher Pollard
arXiv:2604. 01313v2 Announce Type: replace Abstract: High-fidelity simulations and complex inverse problems, such as detector modeling and unfolding, are computationally intensive bottlenecks across subatomic physics, yet essential for accurate physical interpretation.
By Zeyu Xia, Tyler Kim, Trevor Reed, Judy Fox, Geoffrey Fox, Adam Szczepaniak
arXiv:2512. 07420v3 Announce Type: replace-cross Abstract: Jet identification plays a central role in analyzing data from high-energy collider experiments.
By Md Raqibul Islam, Adrita Khan, Mir Sazzat Hossain, Choudhury Ben Yamin Siddiqui, Md. Zakir Hossan, Tanjib Khan, M. Arshad Momen, Amin Ahsan Ali, AKM Mahbubur Rahman
arXiv:2607. 16144v1 Announce Type: cross Abstract: In this work we demonstrate that a single transformer-based generative model can capture Standard Model structure spanning five decades of invariant mass, from the sub-GeV regime to the TeV continuum, a range that no single Monte Carlo sample covers.
By Midori Kato, Kevin A. Urqu\'ia-Calder\'on, Inar Timiryasov, Oleg Ruchayskiy
Panda Diplomacy introduces a point‑cloud self‑distillation framework that enables a single foundation‑model architecture and objective to be pre‑trained across three distinct particle‑detector modalities—liquid argon time‑projection chambers, collider TPCs, and water Cherenkov detectors—without extensive modification. Using only 1,000 labeled images for downstream adaptation, the resulting Panda V2 model matches or surpasses specialized baselines that require orders of magnitude more supervision, achieving state‑of‑the‑art particle‑clustering performance with 70× fewer labeled events on sPHENIX and up to 1,000× fewer labels on LArTPC data. Linear probes further demonstrate that the model’s latent space captures physically meaningful structures such as particle causality and track curvature.
By Samuel Young, C\'esar Jes\'us-Valls, Kazuhiro Terao