arXiv Machine Learning

When Accuracy Gaps Fail to Certify: Auditing Cross-Domain Recalibration of LLM Judges

The study examines whether accuracy gaps between source and target tasks can certify the failure of scalar recalibration maps for large language model judges. Across thirteen judges, two generators, eight domains, and 1,176 transfers, the accuracy gap only provides a weak lower bound on target calibration error and can predict opposite outcomes. Even with a finite‑sample lower certificate, the method shows low power (0.13) and does not reliably indicate when recalibration will fail.

arXiv Machine Learning
Aug 20

ProxyGuard: Direct Reliability Inference for Randomized Data Release Mechanisms with Shared Targets

ProxyGuard is a new method for assessing the reliability of randomized data release mechanisms that use shared target sets. It offers two modes: named-release mode, which corrects for multiplicity and certifies specific releases, and direct shared-target mode, which evaluates independent mechanism draws on a common target, providing finite-sample reliability guarantees without needing independent target batches. In a registered study, direct mode increased power from 5.6% to 64.2% at a 0.95 reliability level, while named mode performed better under high-signal evidence.

By Dipesh Tharu Mahato, Pramod Dhungana
arXiv AI
Sep 4

Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints

The paper reports a preregistered audit of language‑model judges used as measurement instruments, revealing that the assumption that a model’s responses remain stable over time is invalid. Across nearly 53,000 audited requests, repeat rankings and byte‑identical replays fell far below required reliability thresholds, with three identified mechanisms—label‑to‑meaning bias, candidate gaps below the noise floor, and input permutation noise—explaining the discrepancy. The study proposes a three‑level snapshot‑identity framework, eight design rules, and a reporting checklist to prevent such reliability failures in future evaluations.

By Haoyaun Zhu, Jie Zhang
Hugging Face Trending Papers
Sep 3

Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints

The paper investigates the reliability of language‑model judges used as measurement instruments on shared endpoints. Through two preregistered audits of 52,988 requests, the authors found that repeat rankings and byte‑identical replays fell far short of required thresholds, revealing significant instability. They identify three mechanisms—label‑to‑meaning bias, candidate gaps below the noise floor, and input permutation noise—that explain the gap, and propose a snapshot‑identity ladder, design rules, and a reporting checklist to mitigate such failures.