arXiv Machine Learning

When 10,000 Windows Are Not 10,000 Tests: Auditing Statistical Confidence in Sliding-Window Time-Series Classification

The paper critiques the common practice of evaluating sliding‑window time‑series classifiers on thousands of overlapping test windows, noting that such windows are not independent. It introduces an audit framework that maps performance claims to specific aggregation rules and dependence‑robust inference methods, demonstrating that high overlap inflates Type‑I error and variance estimates. Empirical audits on WISDM and HARTH datasets reveal that increasing test rows yields only modest gains in independent information, and that accuracy‑difference intervals widen with overlap, while Macro‑F1 may favor certain models.

arXiv AI
Sep 18

Correct Now, Insufficient Later: Auditing Update Sufficiency in Context Compression

The paper investigates how memory systems can answer a current query correctly yet fail to retain distinctions needed for later updates. Using a paired‑history audit, the authors evaluate 24 history pairs across six synthetic mechanisms and two model backends, achieving perfect reveal accuracy on DeepSeek and high accuracy on GLM. Record‑level audits reveal specific failures in structured reveal memories and frontier late‑reference adequacy, and the authors test a label‑equivariant repair that only partially restores correctness.

By Guangzhe Zhang
arXiv Machine Learning
Jul 9

Prior-matched evaluation of operational Earth-observation classifiers: a three-number reporting method demonstrated on Sentinel-1 internal-wave detection

arXiv:2607. 07146v1 Announce Type: new Abstract: The Internal Waves Service screens the Sentinel-1 Wave-mode archive for internal solitary waves, routing detections to experts whose adjudication time is the resource the effort exists to conserve.

By Joao Pinelo, Joao Goncalves, Arun Shukla, Adriana Santos-Ferreira
arXiv AI
Sep 3

The Ceiling Is in the Channel: Auditing Learner Gaps and Measurement Frontiers in Clinical Prediction

The paper introduces a framework that distinguishes two causes of saturation in clinical prediction: a learner gap, where the model fails to use available information, and a measurement‑channel ceiling, where the recorded variables limit performance. It provides theoretical characterizations, finite‑sample diagnostics, and empirical audits across three large cohorts, showing that well‑tuned models approach the frontier while deficient learners leave large gaps. A PRISMA‑guided synthesis across 104 tasks reveals consistent channel‑level patterns, suggesting that improving the learner or the measurement channel can audit and potentially lift performance.

By Sayeed Shafayet Chowdhury, Nusrat Jahan, Snehasis Mukhopadhyay, Shiaofen Fang, Vijay R. Ramakrishnan
arXiv AI
Sep 23

TimeLitmus: A Diagnostic Benchmark for Cross-Modal Understanding and Explanation Faithfulness in Event-Conditioned Time-Series Prediction

arXiv:2609.24677v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to make predictions from numerical time-series histories and textual events. Yet accuracy alone cann...

By Jie Gong, Maowei Jiang, Zhiwei Liu, Yankai Chen, Guojun Xiong, Xue Liu, Min Peng, Qianqian Xie, Sophia Ananiadou
arXiv AI
Sep 25

How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure

The paper investigates the reliability of ranking tables produced by small-sample evaluations of large language models (LLMs). Using LLM‑inferred prompt structure across eight model variants, the authors find that prompt‑structure recovery is highly unstable, with only the bottom of the ranking consistently reproducible. They demonstrate that standard evaluation practices can misrepresent model performance and propose reporting practices to improve transparency.

By Dipankar Sarkar