Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

18,285 stories · RSS feed

arXiv Machine Learning
Jul 2

TiRex-2: Generalizing TiRex to Multivariate Data and Streaming

arXiv:2607. 01204v1 Announce Type: new Abstract: We introduce TiRex-2, a recurrent xLSTM-based time series foundation model that generalizes the univariate TiRex to multivariate forecasting with both past and future covariates.

By Patrick Podest, Marco Pichler, Elias B\"urger, Levente Z\'olyomi, Bernhard Voggenberger, Wilhelm Berghammer, Daniel Klotz, Sebastian B\"ock, G\"unter Klambauer, Sepp Hochreiter
arXiv AI
Jul 2

PaAno: Patch-Based Representation Learning for Time-Series Anomaly Detection

arXiv:2602. 01359v3 Announce Type: replace-cross Abstract: Although recent studies on time-series anomaly detection have increasingly adopted ever-larger neural network architectures such as transformers and foundation models, they incur high computational costs and memory usage, making them impractical for real-time and resource-constrained scenarios.

By Jinju Park, Seokho Kang
arXiv Machine Learning
Jul 2

Explainability in mulimodal deep transformation models for stroke outcome prediction

arXiv:2504. 06299v2 Announce Type: replace-cross Abstract: Multimodal prediction models based on imaging and clinical data are increasingly used for clinical decision support, yet their interpretability remains limited.

By Lisa Herzog, Jonas Br\"andli, Maurice Schneeberger, Loran Avci, Nordin Dari, Martin H\"ansel, Hakim Baazaoui, Pascal B\"uhler, Susanne Wegener, Beate Sick