Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores
arXiv:2608. 02985v1 Announce Type: new Abstract: The standard check for contamination in LLM backtests is simple: compare scores before and after the training cutoff.
arXiv:2604. 04199v2 Announce Type: replace Abstract: Twenty-eight within-subject counterfactual experiments across 2,047 iid tabular datasets, plus a boundary experiment on 129 temporal datasets, measure the severity of four data leakage classes in machine learning.
arXiv:2608. 02985v1 Announce Type: new Abstract: The standard check for contamination in LLM backtests is simple: compare scores before and after the training cutoff.
arXiv:2606. 11267v1 Announce Type: new Abstract: Data leakage -- contamination of a model with information unavailable at baseline -- is the dominant reproducibility failure in machine-learning-based science, yet detection tools require training code, external data, or domain expertise.
The paper introduces a new taxonomy for benchmark contamination that categorizes leakage by the mitigation it defeats—direct, derivative, temporal, distributional, and acquired—covering both training‑time and evaluation‑time scenarios. It proposes a four‑field disclosure protocol to record contamination status alongside benchmark scores, and provides a JSON schema, validator, and examples. An empirical study of 41 documents using a pre‑registered instrument shows limited reporting of contamination types and variable reliability, highlighting gaps in current disclosure practices.
arXiv:2608. 00144v1 Announce Type: new Abstract: Membership inference (MIA) on language models is usually summarised by an aggregate ROC-AUC, but such evaluations are confounded: model-free blind baselines separate members from non-members from surface text alone.
LeakScale is an interventional framework that estimates the causal effect of benchmark exposure on model performance. By creating new executable tasks that require private, family‑specific information not present in the public task, LeakScale controls access to that information and measures the resulting change in accuracy. Across 2,048 families, two model families, two domains, and 262,144 generations, exposure consistently improves accuracy, with gains ranging from +7.17 to +27.31 percentage points, thereby distinguishing whether benchmark contact occurred from how strongly a reported score depends on it.
arXiv:2607. 16811v4 Announce Type: replace Abstract: Drift detectors that work tend not to explain themselves, and drift detectors that explain themselves tend to fail in high dimension.
The paper critiques the common practice of evaluating sliding‑window time‑series classifiers on thousands of overlapping test windows, noting that such windows are not independent. It introduces an audit framework that maps performance claims to specific aggregation rules and dependence‑robust inference methods, demonstrating that high overlap inflates Type‑I error and variance estimates. Empirical audits on WISDM and HARTH datasets reveal that increasing test rows yields only modest gains in independent information, and that accuracy‑difference intervals widen with overlap, while Macro‑F1 may favor certain models.
arXiv:2607. 01025v1 Announce Type: cross Abstract: Radio-frequency (RF) sensing is a central modality for counter-unmanned-aerial-system (counter-UAS) defence because it exploits the control, telemetry, and video links between a drone and its operator.
arXiv:2506. 20893v5 Announce Type: replace-cross Abstract: In this paper, we reveal a significant shortcoming in class unlearning evaluations: overlooking the underlying class geometry can cause information leakage about the forgotten class.
The paper introduces LeakGauge, a method that appends a suffix to a model’s input to gauge the risk of context leakage before decoding. By mapping prefill token probabilities to an attack‑risk score, LeakGauge achieves high AUROC (0.944–0.996) across 11 large language models, including GLM‑5.2 and Kimi‑K3, and remains robust to language changes and different attack styles. The approach also demonstrates sensitivity to internal leakage directions and can be implemented with fewer than 0.5K additional parameters and minimal latency.
arXiv:2608. 00144v2 Announce Type: replace Abstract: Membership inference (MIA) on language models is usually summarised by aggregate ROC-AUC, but such evaluations are confounded: model-free blind baselines can separate members from non-members using surface text alone.
The paper introduces a subgroup-targeted membership inference game to audit differentially private synthetic text releases, revealing that existing average-case attacks miss significant leakage to vulnerable subgroups. An extensive audit across 32 proxies, four datasets, three generation methods, and five privacy budgets shows that DP reduces overall leakage but leaves concentrated, uneven residual risk, especially for high-risk records. The study demonstrates that which records leak is determined by the release mechanism rather than the records themselves, challenging record-level risk assessment.