When 10,000 Windows Are Not 10,000 Tests: Auditing Statistical Confidence in Sliding-Window Time-Series Classification
Read the original on arXiv Machine Learning →The paper critiques the common practice of evaluating sliding‑window time‑series classifiers on thousands of overlapping test windows, noting that such windows are not independent. It introduces an audit framework that maps performance claims to specific aggregation rules and dependence‑robust inference methods, demonstrating that high overlap inflates Type‑I error and variance estimates. Empirical audits on WISDM and HARTH datasets reveal that increasing test rows yields only modest gains in independent information, and that accuracy‑difference intervals widen with overlap, while Macro‑F1 may favor certain models.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.