arXiv:2608. 16190v1 Announce Type: cross Abstract: Trusted monitoring has a cheap, trusted model score a stronger untrusted model's actions, and a diverse ensemble of them beats a single stronger monitor at matched cost.
By Anik Jha
The paper introduces the twin‑prefix framework to evaluate how the size of the verification unit—i.e., how many actions a pre‑execution LLM monitor reviews in one call—affects its performance. By pairing each gold plan with a twin that differs by a single write and injecting a controlled error, the authors isolate the impact of review length on catch rates and false rejections. Their findings show that longer review windows increase rejection rates but do not improve discrimination, with the highest informedness occurring at one or two actions across all judges and domains.
By Yuchen Han, Cheng Yan, Wuyang Zhang
The paper demonstrates that aggregate accuracy figures for chain‑of‑thought (CoT) monitors can be misleading because a large portion of detected hacks rely solely on action patterns rather than reasoning. By rewriting only the agent’s reasoning to appear truthful while keeping actions identical, the authors show that the monitor’s performance on the reasoning‑dependent subset collapses dramatically, yet the overall pooled accuracy drops only modestly. The study reveals that CoT monitors are fragile when reasoning is the key signal and that accuracy should be reported separately for this subset.
By Shikhar Shiromani, Leo Richter
arXiv:2607. 13075v1 Announce Type: cross Abstract: Context can change whether a request is harmful without changing its topic or surface form.
By Dominik Schwarz
The paper introduces HackProbe, a black‑box monitoring tool that can be attached to any self‑evolving language model loop without accessing internal weights or activations. HackProbe uses a fixed‑distribution comparison core and a rotated fresh layer to detect reward hacking through four statistical tests, and it can immunize the model by selecting honest candidates from the proposal pool. Experiments on a controlled host with injected hacking channels show that HackProbe achieves higher AUROC and lower false‑positive rates than the strongest baseline, and its bandwidth‑limited reselection improves true capability under hacking more than it harms clean runs.
By Rongxin Yang, Yang Liu, Shang Luo, Haoxuan Jia, Chongyang Zhang, Hao Zheng, Yingguang Yang, Yulin Huang, Jianshen Zhang, Yongzhi Qi, Kefu Xu, Congjing Ran, Bin Chong
arXiv:2606. 10456v1 Announce Type: cross Abstract: AI-control monitors score individual agent actions to detect misbehavior, but real harm can be distributed across many benign-looking steps, each individually below any per-step alarm.
By Zhang Qinqin, Gao Yuze