arXiv Machine Learning

Pre-Training for Simulation-Based Science: A Study on Jet Foundation Model Training Objectives

arXiv:2606. 14870v1 Announce Type: cross Abstract: Foundation models (FMs) trained on large datasets and fine-tuned on downstream tasks have emerged as a powerful paradigm in AI for science.

arXiv Machine Learning
Aug 27

Cross-Domain Transfer with Particle Physics Foundation Models: From Jets to Neutrino Interactions

The paper investigates whether two particle physics foundation models, OmniLearned and ParticleViT, pretrained on high‑energy proton–proton and electron–proton collisions can transfer knowledge to a low‑energy neutrino experiment. Using MINERvA neutrino–nucleus scattering data, the authors evaluate the models on energy regression and charged‑current pion classification tasks, finding that the pretrained models outperform similarly sized models trained from scratch, with OmniLearned excelling in regression and ParticleViT in classification. When the same transformer architecture is initialized from unrelated text pretraining (BERT), the performance advantage is minimal for classification and absent for regression, indicating that particle‑level foundation models capture inductive biases that generalize across energy scales, detector technologies, and physics processes.

By Gregor Krzmanc, Vinicius Mikuni, Benjamin Nachman, Callum Wilkinson
arXiv Machine Learning
Sep 10

Mind the Gap: Navigating Inference with Optimal Transport Maps

The paper introduces a model calibration method using optimal transport to address discrepancies between simulation and experimental data in high-dimensional machine learning applications. Applied to jet tagging in particle physics, the technique calibrates a 128‑dimensional latent representation from a general‑purpose classifier, ensuring downstream derived quantities are properly calibrated. This enables more reliable use of foundation models for jet flavor analysis in LHC experiments and offers a general framework for correcting high‑dimensional simulations across scientific fields.

By Malte Algren, Tobias Golling, Francesco Armando Di Bello, Christopher Pollard
arXiv Machine Learning
Jun 5

PI-JEPA: Label-Free Surrogate Pretraining for Coupled Multiphysics Simulation via Operator-Split Latent Prediction

arXiv:2604. 01349v4 Announce Type: replace Abstract: Reservoir simulation workflows face a fundamental data asymmetry: input parameter fields (geostatistical permeability realizations, porosity distributions) are free to generate in arbitrary quantities, yet existing neural operator surrogates require large corpora of expensive labeled simulation trajectories and cannot exploit this unlabeled structure.

By Brandon Yee, Pairie Koh
arXiv AI
Jun 16

JetParticle-JEPA: An Efficient Self-Supervised Representation Learning method for Jet Tagging in High-Energy Physics

arXiv:2606. 14813v1 Announce Type: cross Abstract: Jet tagging at the Large Hadron Collider increasingly relies on deep learning models trained on massive simulated datasets, leading to high computational costs and limited robustness to detector mismodeling.

By Guillaume Letellier (LPCC), Antonin Vacheret (LPCC), Fr\'ed\'eric Jurie
arXiv Machine Learning
Jul 24

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators

arXiv:2607. 20465v1 Announce Type: new Abstract: The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how well LLMs, agents, and data-centric workflows actually prepare training data end to end.

By Hao Liang, Qifeng Cai, Yibo Lin, Jianzhuo Du, Qifeng Xia, Sizhe Qiu, Linzhuang Sun, Meiyi Qiang, Zhaoyang Han, Xiaochen Ma, Bohan Zeng, Ruichuan An, Conghui He, Wentao Zhang
arXiv AI
Sep 15

Generalization Can Emerge in Tabular Foundation Models From a Single Table

The paper demonstrates that a tabular foundation model can achieve strong generalization using only a single real table for self‑supervised pre‑training, challenging the belief that large synthetic or real datasets are necessary. By systematically pre‑training and evaluating across diverse benchmarks, the authors show that the number and quality of tasks that can be derived from a dataset are critical for downstream performance. This finding suggests that carefully constructed task sets from limited data can enable effective transfer learning in tabular models.

By Junwei Ma, Nour Shaheen, Alex Labach, Amine Mhedhbi, Frank Hutter, Anthony L. Caterini, Valentin Thomas