arXiv:2601.11716v2 Announce Type: replace-cross
Abstract: Accurate and efficient detector simulation is essential for modern collider experiments. To reduce the high computational cost, various fast...
By Thorsten Buss, Henry Day-Hall, Frank Gaede, Gregor Kasieczka, Katja Kr\"uger
arXiv:2606. 16035v1 Announce Type: cross Abstract: Modern particles physics experiments have demonstrated an increasing need for fast, high-fidelity detector simulation as detector components have improved and subsequent computational requirements approach the limits of available resources.
By Cole Granger, James Giroux, Richard Tyson, Maurizio Ungaro, Cristiano Fanelli
arXiv:2606. 04165v1 Announce Type: cross Abstract: High-precision calorimeter simulation at current and future colliders imposes rapidly growing computational demands, motivating the development of machine-learning surrogates for traditional Monte Carlo tools such as Geant4.
By Cheng Jiang, Sitian Qian, Kevin Pedro, Oz Amram, Huilin Qu, Maggie Voetberg
The paper introduces a particle‑level generative model that uses residual‑quantized full‑event data to enable fast, ML‑based surrogate simulation for collider events. It demonstrates conditional generation from detector‑stable particles, explores scaling across dataset and model sizes, and shows that token‑level loss predicts downstream physical fidelity. The work offers an empirical framework for scalable collider full‑event generation using residual‑quantized representations.
By Dan Godi, Dmitrii Kobylianskii, Eilam Gross
arXiv:2609.23385v1 Announce Type: cross
Abstract: Data acquisition (DAQ) systems at future particle physics experiments stand to benefit from the extremes of AI/ML development: large-scale foundation...
By Gia Ancone, Qibin Liu, Liangyu Wu, Julia Gonski
arXiv:2606. 14813v1 Announce Type: cross Abstract: Jet tagging at the Large Hadron Collider increasingly relies on deep learning models trained on massive simulated datasets, leading to high computational costs and limited robustness to detector mismodeling.
By Guillaume Letellier (LPCC), Antonin Vacheret (LPCC), Fr\'ed\'eric Jurie
arXiv:2512. 04149v2 Announce Type: replace-cross Abstract: Next token prediction is an attractive pre-training task for jet foundation models, in that it is simulation free and enables excellent generative capabilities that can transfer across datasets.
By Joschka Birk, Anna Hallin, Gregor Kasieczka, Nikol Madzharova, Ian Pang, David Shih
arXiv:2606. 14373v1 Announce Type: cross Abstract: The workflow from particle collision to physics analysis passes through a series of reconstruction steps that are traditionally modular and disconnected, with no shared representation linking low-level detector data to high-level analysis tasks.
By Farouk Mokhtar, Joosep Pata, Michael Kagan, Javier Duarte
arXiv:2609.13205v1 Announce Type: cross
Abstract: Sparse long-context inference requires efficient token retrieval in both prefill and decode. Existing methods often use different retrieval strategie...
By Xu Yang, Jiapeng Zhang, Zhangke, Changjian Chen, Yuxin Chen, Feiqiang Sun, Chengguang Xu, Feng Jin, Zhuo Tang
Panda Diplomacy introduces a point‑cloud self‑distillation framework that enables a single foundation‑model architecture and objective to be pre‑trained across three distinct particle‑detector modalities—liquid argon time‑projection chambers, collider TPCs, and water Cherenkov detectors—without extensive modification. Using only 1,000 labeled images for downstream adaptation, the resulting Panda V2 model matches or surpasses specialized baselines that require orders of magnitude more supervision, achieving state‑of‑the‑art particle‑clustering performance with 70× fewer labeled events on sPHENIX and up to 1,000× fewer labels on LArTPC data. Linear probes further demonstrate that the model’s latent space captures physically meaningful structures such as particle causality and track curvature.
By Samuel Young, C\'esar Jes\'us-Valls, Kazuhiro Terao
The paper introduces a data‑driven method for pairing events at the Large Hadron Collider using the energy mover's distance (EMD) to measure similarity, thereby creating augmentation‑free views for self‑supervised pre‑training. By matching distinct events based on EMD, the approach preserves the physics content of each event without handcrafted distortions. Experiments on QCD jets demonstrate that this pairing technique yields semantic jet embeddings with downstream discrimination power comparable to or better than traditional augmentation‑based baselines.
By Ho Fung Tsoi, Dylan Rankin
arXiv:2606. 17500v1 Announce Type: new Abstract: Transformer-based models achieve strong performance for jet tagging at the CERN LHC, but deploying them in low-latency, resource-constrained trigger systems is challenging.
By Gram Koski, Sean Lipps, Zhenghua Ma, G. Abarajithan, Ryan Kastner