arXiv AI By Wojciech Zarzecki, Jan Dubi\'nski, Sebastian Cygert

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection

Read the original on arXiv AI →

arXiv:2606. 03305v1 Announce Type: new Abstract: Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.