arXiv Machine Learning

CleanScore: Black-Box Benchmark Audits with Negative Controls and Sensitivity Bounds

CleanScore is a black‑box audit framework that uses only scored outputs to assess whether models have been exposed to benchmark questions. It creates a public form and two fresh, independently written forms for each question, reports an interval for the public‑form advantage, and employs a private negative‑control bank with a transport radius to separate exposure from normal form mismatch. In a registered audit of five open models on 200 GSM8K and 200 ARC‑Challenge items, CleanScore found no exposure‑consistent advantage, bounding surface‑form inflation below five points, while also demonstrating how leaked items can inflate accuracy on unseen paraphrases and how planted advantages can be partially detected even after rewriting. whyItMatters":"The study shows that CleanScore can detect and quantify exposure effects in benchmark models, providing a more nuanced understanding of model performance beyond simple accuracy scores."

arXiv Computation and Language
Sep 24

Beyond Overlap: Estimating the Causal Effect of Benchmark Exposure

LeakScale is an interventional framework that estimates the causal effect of benchmark exposure on model performance. By creating new executable tasks that require private, family‑specific information and controlling access to that information, LeakScale measures the control‑adjusted change in accuracy. Across 2,048 families, two model families, two domains, and 262,144 generations, exposure consistently improves accuracy by 7.17 to 27.31 percentage points, separating the question of whether contact occurred from how much a score depends on it.

By Divyansh Singh
arXiv Computation and Language
6d ago

LeakScale: Estimating the Causal Effect of Benchmark Exposure

LeakScale is an interventional framework that estimates the causal effect of benchmark exposure on model performance. By creating new executable tasks that require private, family‑specific information not present in the public task, LeakScale controls access to that information and measures the resulting change in accuracy. Across 2,048 families, two model families, two domains, and 262,144 generations, exposure consistently improves accuracy, with gains ranging from +7.17 to +27.31 percentage points, thereby distinguishing whether benchmark contact occurred from how strongly a reported score depends on it.

By Divyansh Singh
arXiv AI
Sep 1

Benchmark Contamination: A Taxonomy Organized by Defeated Mitigation

The paper introduces a new taxonomy for benchmark contamination that categorizes leakage by the mitigation it defeats—direct, derivative, temporal, distributional, and acquired—covering both training‑time and evaluation‑time scenarios. It proposes a four‑field disclosure protocol to record contamination status alongside benchmark scores, and provides a JSON schema, validator, and examples. An empirical study of 41 documents using a pre‑registered instrument shows limited reporting of contamination types and variable reliability, highlighting gaps in current disclosure practices.

By Johanna Angulo, V\'ictor Yeste, Hector Espinos-Morato
arXiv Machine Learning
1d ago

When a Data Artifact Isn't a Shortcut: Causal Auditing of Synthetic RLVR Corpora

The paper audits whether synthetic distractors in RLVR corpora act as shortcuts for learning policies. A classifier using only surface statistics barely outperforms chance, and manual inspection reveals that code distractors are almost identical to correct answers. Experiments with a paraphrase‑matched control show no exploitation advantage for the unmodified data, indicating that the detectable artifact was not used by the policy.

By Esther Xin
arXiv AI
Sep 24

Can LLMs Catch a Rigged Backtest? A Clean-Control Calibration Benchmark

The paper introduces a 96‑item paired benchmark for evaluating large language models (LLMs) on backtest auditing, where each flawed backtest is matched with a clean control that keeps strategy, dates, code style, labels, and reporting scaffold constant while altering a single methodological detail. A deterministic scorer distinguishes flaw recall, clean‑control false positives, evidence localization, and fix relevance. Experiments on 1440 cached audits from four text endpoints show that the DeepSeek auditor achieves perfect closed and clean‑aware code recall, yet open prompts over‑flag 93.8% of clean controls, and clean‑aware specificity is 87.5% even when recall saturates. Introducing a clean‑aware warning eliminates 20.8% false positives to 0% without affecting recall, though the budget anchor still flags many clean controls. Reporting clean‑control rates provides a clearer differentiation among models than reporting recall alone.

By Makar Ulesov, Vladislav Smirnov, Omar Ibrahim, Arsenii Bobovnikov
arXiv AI
Sep 11

Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States

The paper introduces a reference‑based bias detection method that audits hidden‑state representations of language models by encoding sentences as similarities to a fixed set of anchor sentences. This relative representation allows comparison across model variants, such as before and after fine‑tuning, and yields a metric called Representational Bias Shift (ΔB). ΔB correlates strongly with output‑level bias changes, can detect bias‑increasing checkpoints with high ROC AUC, and is computationally efficient, requiring only a few minutes and far less compute than traditional benchmarks.

By Marek Jeli\'nski, Jan Dubi\'nski, Maciej Chrabaszcz, Sebastian Cygert