Not All Variables Agree: Reliability-Aware Variable-Wise Gradient Surgery for Multivariate Time-Series Forecasting
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2503. 24007v4 Announce Type: replace-cross Abstract: In time series forecasting, covariates represent external factors that influence target variables.
The paper introduces Internal Dual-Wiener routing (Internal‑DW), a backward‑only method that weight‑balances internal gradient routes in autoregressive forecasting. By estimating bounded Wiener gains for identity and nonlinear paths, it suppresses unpredictable noise while preserving predictable learning signals, reducing forecast error by 5.2%–13.8% on four weak‑drive testbeds compared to full BPTT and outperforming gradient clipping, Jacobian regularization, and truncated BPTT in most cases. The approach shows that long‑horizon supervision can be effective without trusting every backward gradient equally.
arXiv:2608. 01857v1 Announce Type: new Abstract: The direction of change --- whether a series will move up or down --- is often as important as its exact value in decisiondriven applications such as risk management and financial forecasting.
arXiv:2608. 05742v1 Announce Type: cross Abstract: Multivariate time series forecasting presents unique challenges because future variables often co-evolve under shared system dynamics.
RATL is a plug‑in method for multivariate time‑series forecasting that uses a frozen base forecaster to build a memory of its historical forecast residuals. During inference, RATL retrieves residual trajectories from similar past contexts and employs a set‑aware router to combine them, providing learned feedback correction. Experiments demonstrate that this residual‑retrieval approach improves the performance of the base forecaster across various benchmarks and backbones.
arXiv:2601. 22631v2 Announce Type: replace-cross Abstract: The application of data-driven remaining useful life (RUL) prediction has long been constrained by the availability of large amount of degradation data.