arXiv Machine Learning By Lucas Pinto

Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages

Read the original on arXiv Machine Learning →

arXiv:2607. 06596v1 Announce Type: cross Abstract: Trusted monitoring is a central defense in AI control: a cheaper trusted model scores an untrusted model's actions for sabotage, and the most suspicious are audited or deferred.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.