Learning Where to Simulate: Generative Active Sampling for Online PDE Surrogate Training
arXiv:2606. 09949v1 Announce Type: cross Abstract: Data-driven PDE surrogates are trained with data produced by numerical PDE solvers.
arXiv:2608. 06107v1 Announce Type: new Abstract: Machine learning offers a promising avenue to accelerate physical simulations by replacing computationally expensive traditional Partial Differential Equation (PDE) solvers with fast, differentiable surrogate models.
arXiv:2606. 09949v1 Announce Type: cross Abstract: Data-driven PDE surrogates are trained with data produced by numerical PDE solvers.
arXiv:2607. 00196v1 Announce Type: new Abstract: Many scientific systems exhibit uncertainty from stochastic forcing, unresolved degrees of freedom, or imperfect observations, making reliable surrogate forecasting fundamentally distributional rather than pointwise.
SimCast‑S2S is a generative latent‑diffusion model designed for probabilistic subseasonal‑to‑seasonal precipitation forecasting. It tackles three key challenges: it uses a diffusion pipeline to capture uncertainty, operates in a compact latent space to enable efficient large‑ensemble generation, and leverages transfer learning with low‑rank adaptation to train on limited reanalysis data after pretraining on climate simulations. The model outperforms deep‑learning baselines and competes with, or surpasses, operational systems such as the ECMWF‑S2S baseline without requiring extensive post‑processing.
SimCast‑S2S is a generative latent‑diffusion framework designed for probabilistic subseasonal‑to‑seasonal precipitation forecasting. It tackles three key challenges: it uses a diffusion‑based generative pipeline for uncertainty quantification, operates in a compact latent space learned by VAEs for efficient large‑ensemble generation, and employs transfer learning with LoRA to overcome limited training data. On reanalysis data, it outperforms deep‑learning baselines and competes with or surpasses state‑of‑the‑art operational systems such as ECMWF‑S2S.
arXiv:2605. 05540v2 Announce Type: replace Abstract: Fast surrogate modeling for high-dimensional physical dynamics requires more than low short-term error: useful models must roll out efficiently while preserving the statistical structure of long trajectories.
The paper investigates why latent neural surrogate solvers, which compress physical system dynamics into a lower‑dimensional space, often fail during long‑horizon autoregressive rollouts. It demonstrates that training the latent representation only for reconstruction leads to instability, and proposes a set of training interventions—Koopman operator learning, Hamming noise injection, and multi‑step rollout fine‑tuning—that align the latent space with long‑horizon forecasting. These interventions reduce long‑rollout error by about 40 % and achieve accuracy comparable to full‑resolution models while using far fewer floating‑point operations and GPU memory, enabling stable extrapolation in mesoscale crystal‑plasticity simulations of high‑cycle fatigue.
arXiv:2606. 05264v1 Announce Type: new Abstract: Training robust multivariate time series forecasting models requires large, diverse corpora, yet many real-world domains provide only a handful of observed sequences.
arXiv:2602. 12706v2 Announce Type: replace Abstract: Neural operators have emerged as fast surrogate solvers for parametric partial differential equations (PDEs).
arXiv:2502. 04646v2 Announce Type: replace-cross Abstract: Weighted sampling -- sampling from a probability density function (PDF) proportional to the product of a base PDF and a weight function -- is a fundamental technique with wide-ranging applications in variance reduction, biased sampling, data augmentation, and more.
arXiv:2608. 14744v1 Announce Type: new Abstract: Recovering high-resolution states from sparse, low-resolution observations is a central challenge in scientific machine learning and data assimilation.
arXiv:2606. 27354v1 Announce Type: cross Abstract: Neural surrogate models offer fast approximate mappings from PDE parameters to solutions, but they typically treat solving as a purely statistical task: once trained, they struggle to correct their own constraint violations and extrapolate beyond the training distribution.
arXiv:2606. 27711v1 Announce Type: cross Abstract: We introduce a neural network-based framework for learning time series estimators through a process we term decision-theoretic pretraining.