arXiv AI By V\'ictor Gallego

Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure

Read the original on arXiv AI →

arXiv:2608. 08722v1 Announce Type: cross Abstract: Benchmarks for systems that are optimized against the evaluation signal measure something different from what they claim.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 22

Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles

The paper introduces mutation analysis as a metric for evaluating GPU‑kernel benchmark oracles, injecting over ten thousand faults into verified CUDA implementations of 188 KernelBench problems. It shows that the current official checkers miss 16.9% of faults, with precision faults being especially problematic, and demonstrates that optimized test suites can achieve 98% detection with only two inputs per problem. The study also reveals flaws in existing patches and a fuzzing recipe that incorrectly rejects correct kernels 107 times.

By Mingzhe Du, Anh Tuan Luu, Dong Huang, See-Kiong Ng