arXiv:2607. 29135v1 Announce Type: cross Abstract: Neural operators provide fast surrogates for time-dependent partial differential equations (PDEs) by applying a learned evolution operator recursively to its own predictions, but this autoregressive rollout feeds every prediction error back as input, so local errors accumulate.
By Jiaquan Zhang, Shuxu Chen, Haifan Meng, Yi Lu, Zhihan Lyu, Fan Mo, Wei Dong, Yang Yang, Chaoning Zhang
arXiv:2609.08554v1 Announce Type: new
Abstract: In data-driven training, multivariate time-series forecasting is usually optimized with a scalar loss averaged over samples, variables, and horizons. T...
By Jinwoo Park, Hyeongwon Kang, Pilsung Kang
RATL is a plug‑in method for multivariate time‑series forecasting that uses a frozen base forecaster to build a memory of its historical forecast residuals. During inference, RATL retrieves residual trajectories from similar past contexts and employs a set‑aware router to combine them, providing learned feedback correction. Experiments demonstrate that this residual‑retrieval approach improves the performance of the base forecaster across various benchmarks and backbones.
By Yuchen He, Yueyang Cang, Zhiyuan Ning, Ningyu Wang, Li Shi
arXiv:2608. 07420v1 Announce Type: new Abstract: World models are expected to support imagination over extended temporal horizons, yet most are still trained through local few-step prediction objectives and deployed by recursively rolling out their own predictions.
By Xinyi Li, Zaishuo Xia, Chenjie Hao, Yubei Chen
arXiv:2602. 16224v2 Announce Type: replace Abstract: Time series data are prone to noise in various domains, and training samples may contain low-predictability patterns that deviate from the normal data distribution, leading to training instability or convergence to poor local minima.
By Xu Zhang, Peng Wang, Yichen Li, Wei Wang
arXiv:2609.13840v1 Announce Type: new
Abstract: A contract-logistics spare-parts operator is paid on order-level service: an order counts only if every requested line is fulfilled, yet forecasters ar...
By Joo Ern Chin, Shih-Fen Cheng, Aldy Gunawan
arXiv:2609.09586v1 Announce Type: new
Abstract: Time series foundation models (TSFMs) are increasingly pre-trained on synthetically generated time series trajectories, where the data generating proce...
By Niloy Biswas, Noureddine El Karoui
arXiv:2608.24593v1 Announce Type: new
Abstract: Adaptive optimizers retain gradient history in moment variables, allowing a local change in loss weighting to alter later updates. We examine whether t...
By Jinhui Guo
arXiv:2606. 18650v1 Announce Type: new Abstract: As Large Language Model (LLM) datasets scale to trillions of tokens, data selection has emerged as a critical frontier to filter out uninformative noise and construct adaptive learning trajectories.
By Jiaxing Wang, Deping Xiang, Jin Xu, Zirui Liu, Zicheng Zhang, Guoqiang Gong, Jun Fang, Chao Liu, Pengzhang Liu, Tongxuan Liu, Ke Zhang, Qixia Jiang
arXiv:2608. 01130v1 Announce Type: new Abstract: A broad range of models face the mismatch where they are updated through trajectory losses but are evaluated by downstream task reward.
By Yuyang Shen
arXiv:2606. 01081v1 Announce Type: new Abstract: Decision-focused learning (DFL) trains predictive models by optimizing downstream decision quality rather than standalone prediction accuracy.
By Wyame Benslimane, Tinghan Ye, Pascal Van Hentenryck, Paul Grigas
The paper introduces SCROLL, a method for forecasting multiple observables in stochastic dynamical systems by composing each observable’s likelihood into per‑task free‑routed last‑layer beliefs on a shared backbone. This approach learns unit‑dependent loss scaling directly from data, enabling accurate predictive variance estimation without separate tuning. Experiments on the Ornstein–Uhlenbeck process, stochastic Lorenz‑63, and real air‑quality data show that SCROLL recovers analytic kernels, achieves superior negative log‑likelihood on state and regime tasks, and maintains calibration while reducing hyper‑parameter search costs.
By Pavel Prochazka