LeakScale is an interventional framework that estimates the causal effect of benchmark exposure on model performance. By creating new executable tasks that require private, family‑specific information and controlling access to that information, LeakScale measures the control‑adjusted change in accuracy. Across 2,048 families, two model families, two domains, and 262,144 generations, exposure consistently improves accuracy by 7.17 to 27.31 percentage points, separating the question of whether contact occurred from how much a score depends on it.
By Divyansh Singh
LeakScale is an interventional framework that estimates the causal effect of benchmark exposure on model performance. By creating new executable tasks that require private, family‑specific information not present in the public task, LeakScale controls access to that information and measures the resulting change in accuracy. Across 2,048 families, two model families, two domains, and 262,144 generations, exposure consistently improves accuracy, with gains ranging from +7.17 to +27.31 percentage points, thereby distinguishing whether benchmark contact occurred from how strongly a reported score depends on it.
By Divyansh Singh
The paper introduces a new taxonomy for benchmark contamination that categorizes leakage by the mitigation it defeats—direct, derivative, temporal, distributional, and acquired—covering both training‑time and evaluation‑time scenarios. It proposes a four‑field disclosure protocol to record contamination status alongside benchmark scores, and provides a JSON schema, validator, and examples. An empirical study of 41 documents using a pre‑registered instrument shows limited reporting of contamination types and variable reliability, highlighting gaps in current disclosure practices.
By Johanna Angulo, V\'ictor Yeste, Hector Espinos-Morato
The paper audits whether synthetic distractors in RLVR corpora act as shortcuts for learning policies. A classifier using only surface statistics barely outperforms chance, and manual inspection reveals that code distractors are almost identical to correct answers. Experiments with a paraphrase‑matched control show no exploitation advantage for the unmodified data, indicating that the detectable artifact was not used by the policy.
By Esther Xin
arXiv:2608. 16852v1 Announce Type: new Abstract: Regulatory compliance monitoring in deployed language models is increasingly implemented as a legal and audit control, checking model outputs against written rules spanning data protection, healthcare, financial regulation, and platform policy.
By Saisab Sadhu, Aadit Sengupta, Vinay Kumar Sankarapu, Pratinav Seth
arXiv:2608. 02665v1 Announce Type: cross Abstract: A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form.
By Yongxi Zhou, Junwei Yao, Yuanzhe Liu, Zihan Dong, Wenbo Ye, Jiaxi Wen, Lai Yun Choi