Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores
arXiv:2608. 02985v1 Announce Type: new Abstract: The standard check for contamination in LLM backtests is simple: compare scores before and after the training cutoff.
arXiv:2606. 11267v1 Announce Type: new Abstract: Data leakage -- contamination of a model with information unavailable at baseline -- is the dominant reproducibility failure in machine-learning-based science, yet detection tools require training code, external data, or domain expertise.
arXiv:2608. 02985v1 Announce Type: new Abstract: The standard check for contamination in LLM backtests is simple: compare scores before and after the training cutoff.
arXiv:2606. 03305v1 Announce Type: new Abstract: Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment.
The paper argues that prediction‑based certifications—such as accuracy, calibration, and conformal coverage—are insufficient to guarantee trustworthy AI. It proves a separation theorem showing that a model can appear reliable under all prediction‑side certificates yet differ arbitrarily in explanation fidelity and deployment behaviour. The authors propose a competence envelope framework that combines both prediction and explanation certification to detect such hidden failures.
arXiv:2608. 02786v1 Announce Type: new Abstract: AI systems can fail silently.
The paper introduces a new taxonomy for benchmark contamination that categorizes leakage by the mitigation it defeats—direct, derivative, temporal, distributional, and acquired—covering both training‑time and evaluation‑time scenarios. It proposes a four‑field disclosure protocol to record contamination status alongside benchmark scores, and provides a JSON schema, validator, and examples. An empirical study of 41 documents using a pre‑registered instrument shows limited reporting of contamination types and variable reliability, highlighting gaps in current disclosure practices.
The study audited ten different classifiers—including linear, tree‑ensemble, neural, glass‑box, and tabular foundation models—on national health survey data to predict myocardial infarction. By systematically removing features that could cause target leakage, the authors found that all models’ AUROC scores collapsed into a narrow band, indicating that reported high accuracy in prior work was largely due to leakage rather than model sophistication. The glass‑box explainable boosting machine performed comparably to other models while being much faster, and the authors demonstrated that fairness, calibration, and uncertainty can be audited and repaired without sacrificing performance.
arXiv:2606. 15153v1 Announce Type: new Abstract: Selective prediction with distribution-free risk control promises that, with confidence 1-delta over the calibration draw, the error rate of accepted inputs stays below a user budget alpha.
arXiv:2606. 16110v1 Announce Type: new Abstract: Machine unlearning has been extensively studied in response to growing privacy concerns and regulatory requirements.
arXiv:2608. 00566v1 Announce Type: new Abstract: Post-hoc model explainers such as LIME, SHAP, and Integrated Gradients are widely deployed to audit models in high-stakes sensitive domains, including finance, healthcare, and social welfare.
arXiv:2606. 25004v1 Announce Type: new Abstract: In machine learning, model certification has been identified as an important method for gaining assurance about a model's trustworthiness and quality.
arXiv:2510.24233v2 Announce Type: replace Abstract: Deep generative models are often trained on sensitive data, such as genetic sequences, health data, or more broadly, any copyrighted, licensed or p...
The paper introduces LeakGauge, a method that appends a suffix to a model’s input to gauge the risk of context leakage before decoding. By mapping prefill token probabilities to an attack‑risk score, LeakGauge achieves high AUROC (0.944–0.996) across 11 large language models, including GLM‑5.2 and Kimi‑K3, and remains robust to language changes and different attack styles. The approach also demonstrates sensitivity to internal leakage directions and can be implemented with fewer than 0.5K additional parameters and minimal latency.