arXiv AI By Nelson Gardner-Challis, Jonathan Bostock, Georgiy Kozhevnikov, Morgan Sinclaire, Joan Velja, Alessandro Abate, Charlie Griffin

When can we trust untrusted monitoring? A safety case sketch across collusion strategies

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.