arXiv Machine Learning

Kastor: An efficient fine-tuning strategy for generative emulation of PDE simulations

arXiv:2608. 06107v1 Announce Type: new Abstract: Machine learning offers a promising avenue to accelerate physical simulations by replacing computationally expensive traditional Partial Differential Equation (PDE) solvers with fast, differentiable surrogate models.

arXiv Machine Learning
Jul 2

TRIE: An Evaluation Framework for Stochastic PDE Surrogates

arXiv:2607. 00196v1 Announce Type: new Abstract: Many scientific systems exhibit uncertainty from stochastic forcing, unresolved degrees of freedom, or imperfect observations, making reliable surrogate forecasting fundamentally distributional rather than pointwise.

By Bharat Srikishan, Javier E. Santos, Nikhil Muralidhar, Charles D. Young
arXiv Machine Learning
Aug 28

SimCast-S2S: An Efficient Generative Model for Subseasonal Precipitation Forecasting via Transfer Learning from Climate Simulations

SimCast‑S2S is a generative latent‑diffusion model designed for probabilistic subseasonal‑to‑seasonal precipitation forecasting. It tackles three key challenges: it uses a diffusion pipeline to capture uncertainty, operates in a compact latent space to enable efficient large‑ensemble generation, and leverages transfer learning with low‑rank adaptation to train on limited reanalysis data after pretraining on climate simulations. The model outperforms deep‑learning baselines and competes with, or surpasses, operational systems such as the ECMWF‑S2S baseline without requiring extensive post‑processing.

By Hiep V. Dang, Antonios Mamalakis
arXiv Machine Learning
Sep 4

SimCast-S2S: A Computationally Efficient Diffusion Model for Subseasonal Precipitation Forecasting

SimCast‑S2S is a generative latent‑diffusion framework designed for probabilistic subseasonal‑to‑seasonal precipitation forecasting. It tackles three key challenges: it uses a diffusion‑based generative pipeline for uncertainty quantification, operates in a compact latent space learned by VAEs for efficient large‑ensemble generation, and employs transfer learning with LoRA to overcome limited training data. On reanalysis data, it outperforms deep‑learning baselines and competes with or surpasses state‑of‑the‑art operational systems such as ECMWF‑S2S.

By Hiep V. Dang, Antonios Mamalakis
arXiv Machine Learning
Sep 25

Beyond Compression: Training Latent Representations for Stable Long-Horizon Rollout in Neural Surrogate Solvers

The paper investigates why latent neural surrogate solvers, which compress physical system dynamics into a lower‑dimensional space, often fail during long‑horizon autoregressive rollouts. It demonstrates that training the latent representation only for reconstruction leads to instability, and proposes a set of training interventions—Koopman operator learning, Hamming noise injection, and multi‑step rollout fine‑tuning—that align the latent space with long‑horizon forecasting. These interventions reduce long‑rollout error by about 40 % and achieve accuracy comparable to full‑resolution models while using far fewer floating‑point operations and GPU memory, enabling stable extrapolation in mesoscale crystal‑plasticity simulations of high‑cycle fatigue.

By Andreas E. Robertson, Ashley T. Lenau, John D. Shimanek, Benjamin A. Jasperson, Vivek Oommen, David L. Damm, Krishna Garikipati, Remi Dingreville
arXiv Machine Learning
Jun 5

REGEN: Reference-Guided Synthetic Multivariate Time Series Generation for Forecasting

arXiv:2606. 05264v1 Announce Type: new Abstract: Training robust multivariate time series forecasting models requires large, diverse corpora, yet many real-world domains provide only a handful of observed sequences.

By Moulik Gupta (Birla AI Labs), Dhruv Kumar (Birla AI Labs, Birla Institute of Technology and Science, Pilani), Murari Mandal (Birla AI Labs, Kalinga Institute of Industrial Technology), Saurabh Deshpande (Birla AI Labs)
arXiv AI
Jun 2

Efficient Weighted Sampling via Score-based Generative Models

arXiv:2502. 04646v2 Announce Type: replace-cross Abstract: Weighted sampling -- sampling from a probability density function (PDF) proportional to the product of a base PDF and a weight function -- is a fundamental technique with wide-ranging applications in variance reduction, biased sampling, data augmentation, and more.

By Heasung Kim, Taekyun Lee, Hyeji Kim, Gustavo de Veciana
arXiv AI
Jun 26

Error-Conditioned Neural Solvers

arXiv:2606. 27354v1 Announce Type: cross Abstract: Neural surrogate models offer fast approximate mappings from PDE parameters to solutions, but they typically treat solving as a purely statistical task: once trained, they struggle to correct their own constraint violations and extrapolate beyond the training distribution.

By Haina Jiang, Liam Wang, Peng-Chen Chen, Min Seop Kwak, Seungryong Kim, Brian Bell, Jeong Joon Park