arXiv Machine Learning

Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages

arXiv:2607. 06596v1 Announce Type: cross Abstract: Trusted monitoring is a central defense in AI control: a cheaper trusted model scores an untrusted model's actions for sabotage, and the most suspicious are audited or deferred.