arXiv:2509. 13805v4 Announce Type: replace-cross Abstract: Foundation models have revolutionized natural language processing through a ``train once, deploy anywhere'' paradigm, where a single pre-trained model adapts to countless downstream tasks without retraining.
By Florian Wiesner, Zo\"e J. Gray, Matthias Wessling, Stephen Baek
arXiv:2607. 23377v1 Announce Type: cross Abstract: The largest machine learning models in particle physics are also the most expensive to train, yet the return on scaling a given architecture cannot be estimated before that compute is spent.
By Jan-Lucas Uslu, Benjamin Nachman, Christopher Re
The largest machine learning models in particle physics are also the most expensive to train, yet the return on scaling a given architecture cannot be estimated before that compute is spent. Scaling laws have been fit for jets, but none has yet been shown to predict the performance of models it was not fit on.
arXiv:2503.19081v2 Announce Type: replace
Abstract: Scientific foundation models (SciFMs) aim to learn generalizable representations of physical systems governed by partial differential equations (PD...
By Serge Kotchourko, Amin Totounferoush, Michael W. Mahoney, Steffen Staab
arXiv:2605. 17985v2 Announce Type: replace-cross Abstract: We propose a new method for compressing physics foundation models (PFMs) which is a new trend in AI for Science.
By Chengjie Hong, Feixiang He, Yiheng Zeng, Lulu Kang, He Wang
arXiv:2607. 27501v1 Announce Type: new Abstract: We present a lightweight approach to foundation modeling (\textbf{NEXUS}) that leverages pre-trained learning from collider physics data towards out-of-domain tasks in other scientific datasets, using a fully connected autoencoder model with approximately 3 million parameters.
By Liangyu Wu, Qibin Liu, Alexander Yue, Julia Gonski
arXiv:2508. 12448v2 Announce Type: replace-cross Abstract: In-context learning (ICL) lets large language models (LLMs) solve new tasks from prompts alone, across an ever-widening range of domains, yet the mechanisms underlying this ability remain poorly understood.
By Yeongwoo Song, Jaeyong Bae, Dong-Kyum Kim, Hawoong Jeong
arXiv:2606. 19781v1 Announce Type: cross Abstract: Neural scaling laws describe how model performance improves as a power law in compute, model size, and dataset size.
By Jan-Lucas Uslu, Kevin Greif, Daniel Whiteson, Benjamin Nachman
arXiv:2609.07814v1 Announce Type: new
Abstract: Physics-informed neural networks (PINNs) struggle on PDEs whose governing physics varies across the domain. We trace this to a structural property of s...
By Hanwen Wang, Paris Perdikaris
The paper investigates whether machine learning models can uncover physical patterns in atomistic data without relying on traditional physics-based inductive biases such as geometric locality or graph structures. By training a general-purpose architecture on molecular simulation data, the authors demonstrate that the model autonomously learns interatomic interaction strengths resembling classical electrostatics and identifies interaction cutoffs aligned with established physical models. The study also reports predictable neural scaling behavior and competitive accuracy on certain metrics compared to physics-informed architectures, suggesting that explicit priors may only be necessary when empirically justified.
By Tobias Kreiman, Yutong Bai, Fadi Atieh, Elizabeth Weaver, Eric Qu, Aditi S. Krishnapriyan
The paper introduces a new approach to domain adaptation in physics, addressing the fact that simulations often differ from experimental data not only in nuisances but also in the target quantity distribution. By studying a toy air‑shower benchmark with separate nuisance, simulation, and spectrum shifts, the authors show that standard adversarial adaptation can misalign spectra, leading to bias. They propose adaptive domain adaptation that reweights simulated events to focus on genuine physical mismatches and provide a label‑free rule for selecting the best model configuration.
By Ivan Kharuk (Institute for Nuclear Research of the Russian Academy of Sciences, Moscow Institute of Physics and Technology)
The paper investigates why latent neural surrogate solvers, which compress physical system dynamics into a lower‑dimensional space, often fail during long‑horizon autoregressive rollouts. It demonstrates that training the latent representation only for reconstruction leads to instability, and proposes a set of training interventions—Koopman operator learning, Hamming noise injection, and multi‑step rollout fine‑tuning—that align the latent space with long‑horizon forecasting. These interventions reduce long‑rollout error by about 40 % and achieve accuracy comparable to full‑resolution models while using far fewer floating‑point operations and GPU memory, enabling stable extrapolation in mesoscale crystal‑plasticity simulations of high‑cycle fatigue.
By Andreas E. Robertson, Ashley T. Lenau, John D. Shimanek, Benjamin A. Jasperson, Vivek Oommen, David L. Damm, Krishna Garikipati, Remi Dingreville