Towards Data Science

Why Transformers Need Positional Encoding For Time Series: A Visual Guide

The article explains how transformers, which rely on self‑attention, can lose the natural order of time‑series data when fed scalar observations. It discusses the role of positional encoding in re‑introducing sequence order and provides a visual guide to illustrate this concept.

arXiv Machine Learning
5d ago

Pattern Formation in Transformers

arXiv:2609.37921v1 Announce Type: new Abstract: What are the inductive biases of a Transformer architecture? Existing theory on how the forward pass shapes representations either considers whether Tr...

By Erkan Turan, Gaspard Abel, Maks Ovsjanikov
Towards Data Science
Jul 2

Time-Series LLMs, Explained with t0-alpha

t0-alpha is a decoder-style patch transformer for probabilistic time-series forecasting. Raw series are split into 32-step patches, embedded, processed through causal time-attention and group-attention layers, and decoded into future quantiles rather than a single point forecast.

By Sean Moran
arXiv Computation and Language
Sep 11

Distance generalization in transformers: why bother with positional encoding?

The paper investigates distance generalization in transformer models, focusing on how well they can handle changes in inter-token distances between training and inference while keeping context length fixed. Using two synthetic delay-copy tasks that require copying tokens after finite delays, the authors evaluate the impact of positional encoding schemes (RoPE, ALiBi, and NoPE), the diversity of distances seen during training, and the conditions under which distance transfer learning is beneficial or detrimental. Their comprehensive study highlights the importance of understanding the underlying mechanisms that govern distance generalization in transformers.

By Daniel Henrik Nevermann, Claudius Gros
arXiv Machine Learning
Aug 6

A Mechanistic Analysis of Transformers for Dynamical Systems

arXiv:2512. 21113v2 Announce Type: replace Abstract: Transformers are increasingly adopted for modeling and forecasting time-series, yet their internal mechanisms remain poorly understood from a dynamical systems perspective.

By Gregory Duth\'e, Nikolaos Evangelou, Wei Liu, Ioannis G. Kevrekidis, Eleni Chatzi