arXiv AI By Ruijie Hou, Yueyang Jiao, Zhao Wang, Yingming Li

Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination

Read the original on arXiv AI →

arXiv:2608. 07341v1 Announce Type: cross Abstract: Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.