arXiv AI

Adaptive Patching Is Harder Than It Looks For Time-Series Forecasting

arXiv:2606. 04074v1 Announce Type: cross Abstract: Adaptive patching is a recent and compelling proposal for time-series Transformers: allocate finer patches where the sequence looks locally informative.

arXiv Machine Learning
Sep 22

A Hybrid Attention Model Learning Unified Time-aware Patch Representation for Irregular Multivariate Time Series Forecasting

The paper introduces a hybrid attention model that learns a unified time‑aware patch representation for irregular multivariate time series (IMTS) forecasting. It employs a time‑aware patch encoding to embed variable‑length intra‑patch timestamps, a time bias attention mechanism to adjust for temporal misalignment and asynchronous cross‑channel dependencies, and a hybrid causal mask on a decoder‑only Transformer to balance historical context with autoregressive forecasting. The authors also curate VersaTSA, a 30 B‑observation dataset preserving native sampling sparsity, and demonstrate state‑of‑the‑art zero‑shot performance on three IMTS benchmarks while remaining competitive on regular MTS tasks.

By Zhihao Lin, Li Lin, Qi Zhang, Kaiwen Xia, Shuai Wang, Jialin Qiao
Hugging Face Trending Papers
Jun 25

How Good Can Linear Models Be for Time-Series Forecasting?

Time-series forecasting research has been moving steadily toward larger architectures, from specialized transformers to general-purpose foundation models, on the assumption that capacity is what unlocks accuracy. We take the opposite position: most of the gap can be closed at far lower cost by tuning preprocessing rather than scaling models.

arXiv Machine Learning
Sep 10

PatchFormer: A Patch-Based Time Series Foundation Model with Hierarchical Masked Reconstruction and Cross-Domain Transfer Learning for Zero-Shot Multi-Horizon Forecasting

arXiv:2601.20845v2 Announce Type: replace Abstract: Time series forecasting is a fundamental problem with applications in climate, energy, healthcare, and finance. Many existing approaches require do...

By Olaf Yunus Laitinen Imanov, Derya Umut Kulali, Taner Yilmaz
arXiv AI
2d ago

Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts

The paper introduces LOOM, a method for scaling looped mixture‑of‑experts (MoE) Transformers beyond the typical two‑loop limit. LOOM addresses two key obstacles: it stabilizes deep recurrence by bounding residual variance and re‑injecting the input embedding, and it prevents expert selection collapse by using per‑loop routers and a looping residual to maintain computational diversity. Experiments on 100 M–1.7 B parameter models show stable scaling to 9–12 loops, with significant perplexity reductions and zero‑shot accuracy gains under near‑iso‑FLOP conditions.

By Di He, Pengxiang Li, Da Chang, Qingyan Meng, Lu Yin, Shiwei Liu
arXiv Machine Learning
Aug 31

Learning to Difference: Adaptive Reversible Differencing (AdaRDiff) for Time Series Forecasting

AdaRDiff is a new adaptive reversible differencing technique for time‑series forecasting that learns weighted differencing to remove trend and seasonality, stabilizes residuals for forecasting, and then reconstructs the forecast autoregressively. The method offers a closed‑form convolutional implementation that can be GPU‑parallelized, achieving up to 33.7× speedup over naive recurrence. Experiments on eight diverse benchmarks show state‑of‑the‑art accuracy and significant performance gains when integrated into various backbone models, from linear models to Transformers.

By Morad Laglil, Younes Hlal, Marouane El Hadari, Emilie Devijver, Eric Gaussier
arXiv Machine Learning
Sep 2

When Does Online Adaptation Pay on the Edge? A Leakage-Free Evaluation of Warmup, Learning-Rate Selection, and Resource Trade-offs for Time-Series Forecasting

The paper investigates when online adaptation benefits edge time‑series forecasting under distribution drift, using a leakage‑free streaming protocol on six public multivariate datasets. It shows that the warmup budget for static baselines and the choice of learning rate can bias perceived adaptation gains, and that a validation‑only procedure selecting warmup and optimizer rates yields Adam outperforming SGD with momentum in most settings. The study also examines accuracy versus adaptation‑state memory and per‑update latency for different adaptation strategies, highlighting parameter‑efficient variants that are nondominated on the memory axis.

By Takumi Fujimoto, Hiroaki Nishi
arXiv Computer Vision
Aug 27

SHIFT-LLM: Distribution Shift Correction in Depth-Pruned LLMs

SHIFT-LLM is a training‑free post‑pruning correction framework that inserts a Linear Residual Adapter (LRA) at each depth‑pruned site in large language models. Each LRA preserves the original residual identity while adding a lightweight affine correction calibrated via closed‑form least‑squares regression on a small held‑out set, thereby approximating the hidden state that would have been produced by the removed block. Experiments across multiple model families and benchmarks show that SHIFT‑LLM consistently recovers accuracy lost to depth pruning, achieving gains up to +15.7 points on Llama‑3.1‑8B‑Instruct with only a few hundred calibration samples and no gradient computation.

By Ali Bahri, Hang Li, Hongliang Li, Zhitang Chen