The Price of Correlated Tests: How Strict Should a Model Release Gate Be?
Read the original on arXiv Statistics ML →The Flow has not summarised this story yet — read it at arXiv Statistics ML.
The Flow has not summarised this story yet — read it at arXiv Statistics ML.
arXiv:2609.13714v1 Announce Type: new Abstract: An updated model can improve an aggregate metric while degrading a slice that matters to a downstream user. We study checkpoint selection subject to no...
arXiv:2609.37493v1 Announce Type: cross Abstract: Serving an answer from a large language model requires deciding when to abstain, yet a verifier's ranking accuracy alone does not determine the error...
arXiv:2609.39229v1 Announce Type: cross Abstract: Automatic evaluation of faithfulness increasingly relies on a large language model acting as a judge, yet the most reliable judges are proprietary fr...
The paper evaluates hard‑gate candidacy for validators in a deployed generative‑agent system by measuring how well each validator’s firing separates successful from failed builds. Across 13 validators and thousands of builds, only a few checks show statistically significant separation, while many fail to distinguish or never fire. The study highlights that skipped checks are recorded as passes, limiting detectable failure rates and underscoring the need for clearer evaluation records.
arXiv:2608. 01378v1 Announce Type: new Abstract: Design campaigns in chemistry, materials science, and machine learning share a bottleneck: determining how good a candidate truly is requires an expensive evaluation - an experiment, a first-principles simulation, or a full training run.
arXiv:2608. 02985v1 Announce Type: new Abstract: The standard check for contamination in LLM backtests is simple: compare scores before and after the training cutoff.