arXiv Machine Learning

ProxyGuard: Direct Reliability Inference for Randomized Data Release Mechanisms with Shared Targets

ProxyGuard is a new method for assessing the reliability of randomized data release mechanisms that use shared target sets. It offers two modes: named-release mode, which corrects for multiplicity and certifies specific releases, and direct shared-target mode, which evaluates independent mechanism draws on a common target, providing finite-sample reliability guarantees without needing independent target batches. In a registered study, direct mode increased power from 5.6% to 64.2% at a 0.95 reliability level, while named mode performed better under high-signal evidence.

arXiv Machine Learning
Jul 7

The Moving Target: A Longitudinal Audit of Trustworthiness Drift Across Twelve Checkpoints of Open-Source Chat LLMs

arXiv:2607. 02587v1 Announce Type: cross Abstract: Model cards quote trust-benchmark scores without recording when they were measured, and the same number is routinely carried across successive checkpoints of one release line as if the model behind it had not shifted.

By Zhichao Fan, Yanhang Li, Zexin Zhuang, Xian Sun, Yingshuo Wang
arXiv AI
Sep 1

Benchmark Contamination: A Taxonomy Organized by Defeated Mitigation

The paper introduces a new taxonomy for benchmark contamination that categorizes leakage by the mitigation it defeats—direct, derivative, temporal, distributional, and acquired—covering both training‑time and evaluation‑time scenarios. It proposes a four‑field disclosure protocol to record contamination status alongside benchmark scores, and provides a JSON schema, validator, and examples. An empirical study of 41 documents using a pre‑registered instrument shows limited reporting of contamination types and variable reliability, highlighting gaps in current disclosure practices.

By Johanna Angulo, V\'ictor Yeste, Hector Espinos-Morato
arXiv AI
Aug 5

BulkPR-Bench: Benchmarking Queue-Level Governance of Interacting Pull Requests

arXiv:2608. 02685v1 Announce Type: cross Abstract: Coding-agent benchmarks increasingly cover long-horizon, end-to-end, and interactive development, but typically retain one requested outcome or a fixed change sequence.

By Zetong Xiong, Qiao Zhao, Jun Zhang, Xueying Lyu, Zhi Li, Yixiang Tu, Xiaowen Yang, Yunjie Zhang, Yufeng Wang, Zhe Zhang, Kaize Yu, Hanwen Du, Zhongkai Sun, Zhuoxin Liu, Zekun Lin, Jianwen Yang, Ruining Chen, Ying Zhang, Tingxuan Pan, Ke Chen, Shubin Han, Chuanhao Sun, Yehua Yang