arXiv Machine Learning By Xinze Shi, Litian Zhang, Binrui Shi

When 10,000 Windows Are Not 10,000 Tests: Auditing Statistical Confidence in Sliding-Window Time-Series Classification

Read the original on arXiv Machine Learning →

The paper critiques the common practice of evaluating sliding‑window time‑series classifiers on thousands of overlapping test windows, noting that such windows are not independent. It introduces an audit framework that maps performance claims to specific aggregation rules and dependence‑robust inference methods, demonstrating that high overlap inflates Type‑I error and variance estimates. Empirical audits on WISDM and HARTH datasets reveal that increasing test rows yields only modest gains in independent information, and that accuracy‑difference intervals widen with overlap, while Macro‑F1 may favor certain models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 18

Correct Now, Insufficient Later: Auditing Update Sufficiency in Context Compression

The paper investigates how memory systems can answer a current query correctly yet fail to retain distinctions needed for later updates. Using a paired‑history audit, the authors evaluate 24 history pairs across six synthetic mechanisms and two model backends, achieving perfect reveal accuracy on DeepSeek and high accuracy on GLM. Record‑level audits reveal specific failures in structured reveal memories and frontier late‑reference adequacy, and the authors test a label‑equivariant repair that only partially restores correctness.

By Guangzhe Zhang