LeakScale is an interventional framework that estimates the causal effect of benchmark exposure on model performance. By creating new executable tasks that require private, family‑specific information and controlling access to that information, LeakScale measures the control‑adjusted change in accuracy. Across 2,048 families, two model families, two domains, and 262,144 generations, exposure consistently improves accuracy by 7.17 to 27.31 percentage points, separating the question of whether contact occurred from how much a score depends on it.
By Divyansh Singh
CleanScore is a black‑box audit framework that uses only scored outputs to assess whether models have been exposed to benchmark questions. It creates a public form and two fresh, independently written forms for each question, reports an interval for the public‑form advantage, and employs a private negative‑control bank with a transport radius to separate exposure from normal form mismatch. In a registered audit of five open models on 200 GSM8K and 200 ARC‑Challenge items, CleanScore found no exposure‑consistent advantage, bounding surface‑form inflation below five points, while also demonstrating how leaked items can inflate accuracy on unseen paraphrases and how planted advantages can be partially detected even after rewriting.
whyItMatters":"The study shows that CleanScore can detect and quantify exposure effects in benchmark models, providing a more nuanced understanding of model performance beyond simple accuracy scores."
By Jeffery Opoku, David Banahene
The paper introduces a new taxonomy for benchmark contamination that categorizes leakage by the mitigation it defeats—direct, derivative, temporal, distributional, and acquired—covering both training‑time and evaluation‑time scenarios. It proposes a four‑field disclosure protocol to record contamination status alongside benchmark scores, and provides a JSON schema, validator, and examples. An empirical study of 41 documents using a pre‑registered instrument shows limited reporting of contamination types and variable reliability, highlighting gaps in current disclosure practices.
By Johanna Angulo, V\'ictor Yeste, Hector Espinos-Morato
arXiv:2608. 02985v1 Announce Type: new Abstract: The standard check for contamination in LLM backtests is simple: compare scores before and after the training cutoff.
By Zeyu Zhang, Bradly C. Stadie
arXiv:2606. 11267v1 Announce Type: new Abstract: Data leakage -- contamination of a model with information unavailable at baseline -- is the dominant reproducibility failure in machine-learning-based science, yet detection tools require training code, external data, or domain expertise.
By Laurence A. Jacobs
arXiv:2606. 03305v1 Announce Type: new Abstract: Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment.
By Wojciech Zarzecki, Jan Dubi\'nski, Sebastian Cygert