LeakScale is an interventional framework that estimates the causal effect of benchmark exposure on model performance. By creating new executable tasks that require private, family‑specific information and controlling access to that information, LeakScale measures the control‑adjusted change in accuracy. Across 2,048 families, two model families, two domains, and 262,144 generations, exposure consistently improves accuracy by 7.17 to 27.31 percentage points, separating the question of whether contact occurred from how much a score depends on it.
By Divyansh Singh
CleanScore is a black‑box audit framework that uses only scored outputs to assess whether models have been exposed to benchmark questions. It creates a public form and two fresh, independently written forms for each question, reports an interval for the public‑form advantage, and employs a private negative‑control bank with a transport radius to separate exposure from normal form mismatch. In a registered audit of five open models on 200 GSM8K and 200 ARC‑Challenge items, CleanScore found no exposure‑consistent advantage, bounding surface‑form inflation below five points, while also demonstrating how leaked items can inflate accuracy on unseen paraphrases and how planted advantages can be partially detected even after rewriting.
whyItMatters":"The study shows that CleanScore can detect and quantify exposure effects in benchmark models, providing a more nuanced understanding of model performance beyond simple accuracy scores."
By Jeffery Opoku, David Banahene
The paper introduces a new taxonomy for benchmark contamination that categorizes leakage by the mitigation it defeats—direct, derivative, temporal, distributional, and acquired—covering both training‑time and evaluation‑time scenarios. It proposes a four‑field disclosure protocol to record contamination status alongside benchmark scores, and provides a JSON schema, validator, and examples. An empirical study of 41 documents using a pre‑registered instrument shows limited reporting of contamination types and variable reliability, highlighting gaps in current disclosure practices.
By Johanna Angulo, V\'ictor Yeste, Hector Espinos-Morato
arXiv:2608. 02985v1 Announce Type: new Abstract: The standard check for contamination in LLM backtests is simple: compare scores before and after the training cutoff.
By Zeyu Zhang, Bradly C. Stadie
arXiv:2606. 11267v1 Announce Type: new Abstract: Data leakage -- contamination of a model with information unavailable at baseline -- is the dominant reproducibility failure in machine-learning-based science, yet detection tools require training code, external data, or domain expertise.
By Laurence A. Jacobs
arXiv:2606. 03305v1 Announce Type: new Abstract: Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment.
By Wojciech Zarzecki, Jan Dubi\'nski, Sebastian Cygert
arXiv:2609.15017v1 Announce Type: cross
Abstract: Prompt-injection detectors are typically evaluated using aggregate F1 on in-distribution test data, which offers limited insight into behavior under...
By Yusuf Khalid Shire, Sang-Chul Kim
arXiv:2607. 23514v1 Announce Type: cross Abstract: Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence.
By Haorui He, Xinwen Chen, Dacheng Wen, Reynold Cheng, Francis C. M. Lau, Yupeng Li
The paper introduces ToxicBench, a benchmark designed to evaluate how tool‑augmented data agents handle incorrect tool outputs. By pairing clean and poisoned observations across numerical, label, schema, and retrieval errors, the authors assess both the checking process and the final answer adoption. In a 118‑task GPT evaluation, poisoning reduces task success by 26–39 percentage points, revealing that repeated poisoning leads to wrong-answer adoption even after checking, while ordinary retries help only under one‑shot poisoning. Human annotations on 200 trajectories confirm the scoring system’s reliability, showing 96% agreement with task success and supporting the benefits of retries and audit‑based adoption.
By Zifu Tao, Changqing Yin
The paper introduces a 96‑item paired benchmark for evaluating large language models (LLMs) on backtest auditing, where each flawed backtest is matched with a clean control that keeps strategy, dates, code style, labels, and reporting scaffold constant while altering a single methodological detail. A deterministic scorer distinguishes flaw recall, clean‑control false positives, evidence localization, and fix relevance. Experiments on 1440 cached audits from four text endpoints show that the DeepSeek auditor achieves perfect closed and clean‑aware code recall, yet open prompts over‑flag 93.8% of clean controls, and clean‑aware specificity is 87.5% even when recall saturates. Introducing a clean‑aware warning eliminates 20.8% false positives to 0% without affecting recall, though the budget anchor still flags many clean controls. Reporting clean‑control rates provides a clearer differentiation among models than reporting recall alone.
By Makar Ulesov, Vladislav Smirnov, Omar Ibrahim, Arsenii Bobovnikov
arXiv:2606. 05029v1 Announce Type: new Abstract: Controlled experiments are the backbone of machine learning research, but at the scale of modern foundation models, they have become prohibitively expensive.
By Gunnar K\"onig, Martin Pawelczyk, Ulrike von Luxburg, Sebastian Bordt
arXiv:2607. 20827v1 Announce Type: new Abstract: LLM agents choose tools and arguments from context that mixes user requests, tool outputs, retrieved records, memory, and untrusted text.
By Junchi Liao