arXiv Machine Learning

Which Leakage Types Matter? A Quantitative Landscape Across 2,047 Benchmark Datasets

arXiv:2604. 04199v2 Announce Type: replace Abstract: Twenty-eight within-subject counterfactual experiments across 2,047 iid tabular datasets, plus a boundary experiment on 129 temporal datasets, measure the severity of four data leakage classes in machine learning.

arXiv AI
Sep 1

Benchmark Contamination: A Taxonomy Organized by Defeated Mitigation

The paper introduces a new taxonomy for benchmark contamination that categorizes leakage by the mitigation it defeats—direct, derivative, temporal, distributional, and acquired—covering both training‑time and evaluation‑time scenarios. It proposes a four‑field disclosure protocol to record contamination status alongside benchmark scores, and provides a JSON schema, validator, and examples. An empirical study of 41 documents using a pre‑registered instrument shows limited reporting of contamination types and variable reliability, highlighting gaps in current disclosure practices.

By Johanna Angulo, V\'ictor Yeste, Hector Espinos-Morato
arXiv Computation and Language
6d ago

LeakScale: Estimating the Causal Effect of Benchmark Exposure

LeakScale is an interventional framework that estimates the causal effect of benchmark exposure on model performance. By creating new executable tasks that require private, family‑specific information not present in the public task, LeakScale controls access to that information and measures the resulting change in accuracy. Across 2,048 families, two model families, two domains, and 262,144 generations, exposure consistently improves accuracy, with gains ranging from +7.17 to +27.31 percentage points, thereby distinguishing whether benchmark contact occurred from how strongly a reported score depends on it.

By Divyansh Singh
arXiv Machine Learning
5d ago

When 10,000 Windows Are Not 10,000 Tests: Auditing Statistical Confidence in Sliding-Window Time-Series Classification

The paper critiques the common practice of evaluating sliding‑window time‑series classifiers on thousands of overlapping test windows, noting that such windows are not independent. It introduces an audit framework that maps performance claims to specific aggregation rules and dependence‑robust inference methods, demonstrating that high overlap inflates Type‑I error and variance estimates. Empirical audits on WISDM and HARTH datasets reveal that increasing test rows yields only modest gains in independent information, and that accuracy‑difference intervals widen with overlap, while Macro‑F1 may favor certain models.

By Xinze Shi, Litian Zhang, Binrui Shi
arXiv AI
Aug 19

The Model's Tell: Measuring Context-Leakage Attack Signals with Behavior Gauges

The paper introduces LeakGauge, a method that appends a suffix to a model’s input to gauge the risk of context leakage before decoding. By mapping prefill token probabilities to an attack‑risk score, LeakGauge achieves high AUROC (0.944–0.996) across 11 large language models, including GLM‑5.2 and Kimi‑K3, and remains robust to language changes and different attack styles. The approach also demonstrates sensitivity to internal leakage directions and can be implemented with fewer than 0.5K additional parameters and minimal latency.

By Maosen Zhang, Jianshuo Dong, Boting Lu, Wenyue Li, Xiaoping Zhang, Tianwei Zhang, Jie Zhang, Han Qiu
arXiv AI
Sep 11

Subgroup Membership Inference Audits of Differentially Private Synthetic Text

The paper introduces a subgroup-targeted membership inference game to audit differentially private synthetic text releases, revealing that existing average-case attacks miss significant leakage to vulnerable subgroups. An extensive audit across 32 proxies, four datasets, three generation methods, and five privacy budgets shows that DP reduces overall leakage but leaves concentrated, uneven residual risk, especially for high-risk records. The study demonstrates that which records leak is determined by the release mechanism rather than the records themselves, challenging record-level risk assessment.

By Yidan Sun, Viktor Schlegel, Srinivasan Nandakumar, Siew Kei Lam, Anil Anthony Bharath