arXiv AI By Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma

When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits

Read the original on arXiv AI →

arXiv:2608. 07914v1 Announce Type: new Abstract: Behavioral contamination detectors can return "no evidence" either because a benchmark is clean or because the audit has little power.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.