arXiv Statistics ML

Distillation of Synthetic Data for Time Series Foundation Models

arXiv Machine Learning
Jul 2

TRIE: An Evaluation Framework for Stochastic PDE Surrogates

arXiv:2607. 00196v1 Announce Type: new Abstract: Many scientific systems exhibit uncertainty from stochastic forcing, unresolved degrees of freedom, or imperfect observations, making reliable surrogate forecasting fundamentally distributional rather than pointwise.

By Bharat Srikishan, Javier E. Santos, Nikhil Muralidhar, Charles D. Young
arXiv Machine Learning
3d ago

Large Distant Gradients Need Not Be Reliable: reliability-weighted credit assignment for long-horizon autoregressive forecasting

The paper introduces Internal Dual-Wiener routing (Internal‑DW), a backward‑only method that weight‑balances internal gradient routes in autoregressive forecasting. By estimating bounded Wiener gains for identity and nonlinear paths, it suppresses unpredictable noise while preserving predictable learning signals, reducing forecast error by 5.2%–13.8% on four weak‑drive testbeds compared to full BPTT and outperforming gradient clipping, Jacobian regularization, and truncated BPTT in most cases. The approach shows that long‑horizon supervision can be effective without trusting every backward gradient equally.

By Junhao Zhao, David Michael Simberg, Jacob Kang, Colin Connor Kurniawan, Nan Xu
arXiv Machine Learning
Aug 31

Prequential posteriors

The paper introduces prequential posteriors, a Bayesian approach that uses a predictive‑sequential loss function to update deep generative forecasting models (DGFMs) when new data arrive. By adopting a consistency notion suitable for model misspecification, the authors prove that both the loss minimizer and the posterior concentrate on parameters with optimal predictive performance. Scalable inference is achieved with parallelisable waste‑free sequential Monte Carlo samplers that employ preconditioned gradient kernels, and the method is validated on synthetic and real meteorological time‑series data.

By Shreya Sinha-Roy, Richard G. Everitt, Christian P. Robert, Ritabrata Dutta
arXiv Machine Learning
Aug 27

Forecasting Multiple Observables with SCROLL: Score-Trained Uncertainty for Stochastic Dynamics

The paper introduces SCROLL, a method for forecasting multiple observables in stochastic dynamical systems by composing each observable’s likelihood into per‑task free‑routed last‑layer beliefs on a shared backbone. This approach learns unit‑dependent loss scaling directly from data, enabling accurate predictive variance estimation without separate tuning. Experiments on the Ornstein–Uhlenbeck process, stochastic Lorenz‑63, and real air‑quality data show that SCROLL recovers analytic kernels, achieves superior negative log‑likelihood on state and regime tasks, and maintains calibration while reducing hyper‑parameter search costs.

By Pavel Prochazka
arXiv AI
Jul 23

Post-Training in Time Series Foundation Models: A Unifying Framework

arXiv:2607. 20002v1 Announce Type: cross Abstract: Time series foundation models (TSFMs) have emerged as general-purpose models for time series analysis, but pretraining alone is often insufficient for reliable downstream deployment.

By Shifeng Xie, Ambroise Odonnat, Zehao Xiao, Lei Zan, Malik Tiomoko, Lujia Pan, Themis Palpanas, Boris N. Oreshkin, Chenghao Liu, Keli Zhang
Hugging Face Trending Papers
Jul 22

Post-Training in Time Series Foundation Models: A Unifying Framework

Time series foundation models (TSFMs) have emerged as general-purpose models for time series analysis, but pretraining alone is often insufficient for reliable downstream deployment. Bridging this gap requires further intervention to handle domain shift, task heterogeneity, limited supervision, and computational constraints, which motivates post-training as a broad class of methods to adapt, augment, compose, calibrate, or specialize pretrained TSFMs for downstream tasks.

arXiv Machine Learning
Jun 5

Diffusion Models for Adaptive Sequential Data Generation

arXiv:2606. 06007v1 Announce Type: new Abstract: Generating realistic synthetic sequential data is critical in real-world applications across operations research, finance, healthcare, energy systems, and scientific computing, where time-indexed observations are used for prediction, simulation, risk assessment, and data-driven decision-making.

By Haoyang Cao, Minshuo Chen, Yinbin Han, Renyuan Xu