arXiv:2607. 02587v2 Announce Type: replace-cross Abstract: Trust-benchmark scores reported on a chat-LLM release line are often carried across several checkpoints of the same line, as if the underlying model had not shifted between releases.
By Zhichao Fan, Yanhang Li, Zexin Zhuang, Xian Sun, Yingshuo Wang
arXiv:2609.13714v1 Announce Type: new
Abstract: An updated model can improve an aggregate metric while degrading a slice that matters to a downstream user. We study checkpoint selection subject to no...
By Shengwei Zhang, Tao Wu, Fei Qian
arXiv:2608. 07913v1 Announce Type: cross Abstract: Selective-risk certificates promise that accepted outputs meet a declared error target.
By Sanjeda Akter, Ibne Farabi Shihab, Anuj Sharma
arXiv:2607. 02587v1 Announce Type: cross Abstract: Model cards quote trust-benchmark scores without recording when they were measured, and the same number is routinely carried across successive checkpoints of one release line as if the model behind it had not shifted.
By Zhichao Fan, Yanhang Li, Zexin Zhuang, Xian Sun, Yingshuo Wang
Selective-risk certificates promise that accepted outputs meet a declared error target. We develop Fed-SRC, a score-agnostic certificate for federated, differentially private, adaptively monitored retrieval-augmented generation.
The paper introduces a new taxonomy for benchmark contamination that categorizes leakage by the mitigation it defeats—direct, derivative, temporal, distributional, and acquired—covering both training‑time and evaluation‑time scenarios. It proposes a four‑field disclosure protocol to record contamination status alongside benchmark scores, and provides a JSON schema, validator, and examples. An empirical study of 41 documents using a pre‑registered instrument shows limited reporting of contamination types and variable reliability, highlighting gaps in current disclosure practices.
By Johanna Angulo, V\'ictor Yeste, Hector Espinos-Morato
arXiv:2608. 07914v1 Announce Type: new Abstract: Behavioral contamination detectors can return "no evidence" either because a benchmark is clean or because the audit has little power.
By Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma
arXiv:2607. 22926v1 Announce Type: new Abstract: High-impact generative AI makes catastrophic misuse a lifecycle-control problem, not merely a prompt-filtering problem.
By Mahdi Eslamimehr
arXiv:2609.19072v1 Announce Type: cross
Abstract: Large language models are increasingly used for content moderation, but most evaluations still report aggregate accuracy on individual benchmarks. We...
By Yibo Hu
arXiv:2608. 02685v1 Announce Type: cross Abstract: Coding-agent benchmarks increasingly cover long-horizon, end-to-end, and interactive development, but typically retain one requested outcome or a fixed change sequence.
By Zetong Xiong, Qiao Zhao, Jun Zhang, Xueying Lyu, Zhi Li, Yixiang Tu, Xiaowen Yang, Yunjie Zhang, Yufeng Wang, Zhe Zhang, Kaize Yu, Hanwen Du, Zhongkai Sun, Zhuoxin Liu, Zekun Lin, Jianwen Yang, Ruining Chen, Ying Zhang, Tingxuan Pan, Ke Chen, Shubin Han, Chuanhao Sun, Yehua Yang
arXiv:2606. 26185v1 Announce Type: new Abstract: LLM-as-judge ("grader") components are now standard in evaluation harnesses, including safety evaluations where a pass/fail verdict may gate downstream deployment decisions.
By Hiroki Tamba
arXiv:2508. 04064v2 Announce Type: replace-cross Abstract: Horizontal federated learning (HFL) backdoor audits often summarize model behavior through clean accuracy (CA), mean attack success rate (ASR), or a single known-trigger test.
By Tuan Nguyen, Sze Jue Yang, Khoa D. Doan, Chee Seng Chan, Kok-Seng Wong