arXiv Machine Learning By Jie Peng, Rui Wang, Qiang Wang, Zhewei Wei, Bin Tong, Guan Wang, Bo Zheng

From Leakage to Fidelity: Reliable Benchmarking for Temporal Cascade Prediction

Read the original on arXiv Machine Learning →

The paper critiques current temporal cascade prediction benchmarks for relying on leakage-prone random splits, limited datasets, and unexamined protocol effects. It proposes a fidelity-aware benchmarking suite featuring the Full Temporal protocol, overlap-based leakage diagnostics, and analyses of performance inflation and temporal drift. Additionally, it introduces the Taoke e‑commerce dataset with rich features and purchase conversions, and presents CasTemp as a lightweight reference method for scalable evaluation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 19

LiveHouse-TS: An Open-world Living Benchmark for Time Series Foundation Models

LiveHouse-TS introduces an open‑world living benchmark for Time Series Foundation Models, evaluating them prequentially on real future data rather than static test windows. The benchmark captures continuous performance across seasonal changes, distribution shifts, and unexpected events, providing a more realistic assessment of model robustness. Experiments across 11 domains and 17 datasets show that model rankings can dramatically change under this live protocol.

By Haomin Wen, Ziyu Zhou, Qingxiang Liu, Siru Zhong, Yuxuan Liang
arXiv Machine Learning
Jun 9

Benchmark Datasets for Lead-Lag Forecasting on Social Platforms

arXiv:2511. 03877v2 Announce Type: replace Abstract: Social and collaborative platforms emit multivariate time-series traces in which early interactions -- such as views, likes, or downloads -- are followed, sometimes months or years later, by higher impact like citations, sales, or reviews.

By Kimia Kazemian (Department of Computer Science, Cornell University), Zhenzhen Liu (Department of Computer Science, Cornell University), Yangfanyu Yang (Department of Information Science, Cornell University), Katie Luo (Department of Computer Science, Stanford University), Shuhan Gu (Department of Computer Science, Cornell University), Audrey Du (Department of Computer Science, Cornell University), Xinyu Yang (Department of Information Science, Cornell University), Jack Jansons (Department of Computer Science, Cornell University), Kilian Q. Weinberger (Department of Computer Science, Cornell University), John Thickstun (Department of Computer Science, Cornell University), Yian Yin (Department of Information Science, Cornell University), Sarah Dean (Department of Computer Science, Cornell University)
arXiv AI
Jul 15

Scaling Point-in-Time Language Models

arXiv:2607. 11889v1 Announce Type: cross Abstract: Large language models trained on unrestricted internet corpora inevitably embed information from the future, introducing lookahead bias that compromises the validity of backtests and causal inference in finance and the social sciences.

By Bryan Kelly, Semyon Malamud, Johannes Schwab, Teng Andrea Xu