The paper derives precise cost formulas for self‑calibrating monitors that adjust thresholds online to maintain a specified long‑run false‑alarm rate under arbitrary drift. It shows that the guarantee is an accounting identity, independent of the monitored signal, and provides exact evidence identities for both step and ramp drift scenarios, as well as an exact law for the fluctuation of the certificate’s own alarm rate. Additionally, it proves that any monitor designed to tolerate a drift class is blind to all faults in the difference of that class, identifying the blind set for speed‑bounded drift classes and quantifying power outside this set with a sharp Gaussian projection bound.
By Abdou-Raouf Atarmla
arXiv:2607. 19393v1 Announce Type: cross Abstract: While auditing a perturbation-based OOD detector on a document benchmark, we recorded an AUROC of 0.
By Vishnu Bindu Balachandran
arXiv:2608. 07914v1 Announce Type: new Abstract: Behavioral contamination detectors can return "no evidence" either because a benchmark is clean or because the audit has little power.
By Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma
The paper introduces HackProbe, a black‑box monitoring tool that can be attached to any self‑evolving language model loop without accessing internal weights or activations. HackProbe uses a fixed‑distribution comparison core and a rotated fresh layer to detect reward hacking through four statistical tests, and it can immunize the model by selecting honest candidates from the proposal pool. Experiments on a controlled host with injected hacking channels show that HackProbe achieves higher AUROC and lower false‑positive rates than the strongest baseline, and its bandwidth‑limited reselection improves true capability under hacking more than it harms clean runs.
By Rongxin Yang, Yang Liu, Shang Luo, Haoxuan Jia, Chongyang Zhang, Hao Zheng, Yingguang Yang, Yulin Huang, Jianshen Zhang, Yongzhi Qi, Kefu Xu, Congjing Ran, Bin Chong
arXiv:2609.06036v1 Announce Type: new
Abstract: Proposal-based controllers---learned policies, language-model planners, and other black-box \emph{generators}---are increasingly deployed behind runtim...
By Guangxi Wan, Yongbo Xie, Yuqi Liu, Qingwei Dong, Qingxin Li, Hongfei Bai, Peng Zeng
arXiv:2505. 02299v2 Announce Type: replace-cross Abstract: Machine Learning (ML) models are trained on in-distribution (ID) data but often encounter out-of-distribution (OOD) inputs during deployment---posing serious risks in safety-critical domains.
By Daisuke Yamada, Harit Vishwakarma, Ramya Korlakai Vinayak