arXiv Machine Learning

Learning Standard Model structure from LHC data with Riemannian flow matching

arXiv:2607. 16144v1 Announce Type: cross Abstract: In this work we demonstrate that a single transformer-based generative model can capture Standard Model structure spanning five decades of invariant mass, from the sub-GeV regime to the TeV continuum, a range that no single Monte Carlo sample covers.

arXiv Machine Learning
Sep 10

Mind the Gap: Navigating Inference with Optimal Transport Maps

The paper introduces a model calibration method using optimal transport to address discrepancies between simulation and experimental data in high-dimensional machine learning applications. Applied to jet tagging in particle physics, the technique calibrates a 128‑dimensional latent representation from a general‑purpose classifier, ensuring downstream derived quantities are properly calibrated. This enables more reliable use of foundation models for jet flavor analysis in LHC experiments and offers a general framework for correcting high‑dimensional simulations across scientific fields.

By Malte Algren, Tobias Golling, Francesco Armando Di Bello, Christopher Pollard
arXiv Machine Learning
1d ago

Scaling Collider Event Generation with Residual-Quantized Tokens

The paper introduces a particle‑level generative model that uses residual‑quantized full‑event data to enable fast, ML‑based surrogate simulation for collider events. It demonstrates conditional generation from detector‑stable particles, explores scaling across dataset and model sizes, and shows that token‑level loss predicts downstream physical fidelity. The work offers an empirical framework for scalable collider full‑event generation using residual‑quantized representations.

By Dan Godi, Dmitrii Kobylianskii, Eilam Gross
arXiv Machine Learning
Jun 8

ScatterPrism: convergence for generative simulation and inverse problems in particle and nuclear physics

arXiv:2604. 01313v2 Announce Type: replace Abstract: High-fidelity simulations and complex inverse problems, such as detector modeling and unfolding, are computationally intensive bottlenecks across subatomic physics, yet essential for accurate physical interpretation.

By Zeyu Xia, Tyler Kim, Trevor Reed, Judy Fox, Geoffrey Fox, Adam Szczepaniak
arXiv Machine Learning
Sep 14

Learning the Geometry of Collider Events with Metric-Aware Deep Sets

The paper introduces a Deep Sets surrogate for optimal transport (OT) that respects key metric properties—non-negativity, exchange symmetry, and zero self-distance—while leaving the triangle inequality unconstrained. Applied to the Energy Mover's Distance between collider events, the Metric-Aware Particle Flow Network achieves percent‑level mean absolute percentage error and markedly higher inference throughput compared to other exact and approximate methods. The architectural constraints also dramatically reduce triangle‑inequality violations, improving geometric fidelity across a large set of held‑out event triplets.

By Lauren Hay, Rishabh Jain, Matt LeBlanc, Jennifer Roloff
arXiv Machine Learning
Sep 10

Likelihood-Based Unsupervised Anomaly Detection in CMS Dijet Events

arXiv:2609.06686v1 Announce Type: cross Abstract: We present an unsupervised search for anomalous dijet events in proton--proton collision data using neural spline flow density estimation. A normaliz...

By Bhavishya Chebrolu (VIT-AP University, Amaravati, India), Hitesh Rasineni (VIT-AP University, Amaravati, India), Prajwal Aaryan Immadi (VIT-AP University, Amaravati, India)
arXiv Machine Learning
Sep 14

Fast BIB simulation at a future Muon Collider with generative machine learning

The paper presents the first machine learning models for fast generation of beam‑induced background (BIB) in tracking detectors at a future Muon Collider. Two architectures are explored: a high‑fidelity tabular diffusion model and a faster circular spline flow model. Both produce BIB hits and tracks that closely match full simulation results, achieving over an order of magnitude speed‑up while requiring far less computational resources.

By Radha Mastandrea, Shiyu Peng, Benjamin Rosser, Matt LeBlanc
arXiv AI
Jun 16

JetParticle-JEPA: An Efficient Self-Supervised Representation Learning method for Jet Tagging in High-Energy Physics

arXiv:2606. 14813v1 Announce Type: cross Abstract: Jet tagging at the Large Hadron Collider increasingly relies on deep learning models trained on massive simulated datasets, leading to high computational costs and limited robustness to detector mismodeling.

By Guillaume Letellier (LPCC), Antonin Vacheret (LPCC), Fr\'ed\'eric Jurie
arXiv Machine Learning
Sep 17

Similarity Pairing with Energy Mover's Distance for Self-Supervised Pre-Training at the LHC

The paper introduces a data‑driven method for pairing events at the Large Hadron Collider using the energy mover's distance (EMD) to measure similarity, thereby creating augmentation‑free views for self‑supervised pre‑training. By matching distinct events based on EMD, the approach preserves the physics content of each event without handcrafted distortions. Experiments on QCD jets demonstrate that this pairing technique yields semantic jet embeddings with downstream discrimination power comparable to or better than traditional augmentation‑based baselines.

By Ho Fung Tsoi, Dylan Rankin