arXiv:2608. 12652v1 Announce Type: cross Abstract: Benchmark contamination is diagnosed today with n-gram overlap, with likelihood-based membership inference, or with canary strings, and each needs something usually unavailable: the training corpus, a well-chosen test statistic, or foresight at dataset release.
By Florian Braun
The study analyzes 373,019 judgments from LLM‑scored benchmarks, decomposing variance into system, item, judge, and interaction components via generalizability theory. It finds that with a single judge, generalizability converges to a ceiling determined by the system‑by‑judge variance, which is substantially lower in pairwise preference settings, allowing one judge to suffice. The research also reveals significant biases in presentation order and highlights that many published win‑rate claims fall below the measured floor of the benchmarks.
By Atul Anand
arXiv:2606. 25097v1 Announce Type: new Abstract: Speculative decoding accelerates inference by letting a draft model propose tokens for a target model to verify, raising a concrete safety question: at temperature zero, can draft-side behavior leak into safety-scored outputs?
By Sahil Kadadekar
arXiv:2509. 11208v3 Announce Type: replace-cross Abstract: Transformers used for evidence-grounded binary adjudication (e.
By Leon Chlon, Ahmed Karim, Maggie Chlon, MarcAntonio Awada
arXiv:2607. 10202v2 Announce Type: replace Abstract: Forced-choice probes with counterbalanced orientations are a standard tool for measuring language-model "value dispositions," and a concentration/extremity index over repeated draws is read as how sharply a model commits.
By Hong-In Won, Jinseok Jang, Hyoseop Kim
arXiv:2606. 15610v1 Announce Type: cross Abstract: LLM-as-a-judge systems are now routinely used for open-ended model evaluation, where human preference annotation is costly, slow, and difficult to reproduce.
By Hiroyasu Usami, Keisuke Hara, Ayato Tsuboi, Naohiko Matsuda