arXiv AI

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

arXiv:2607. 19321v1 Announce Type: new Abstract: As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted.

arXiv AI
Jul 17

Democratizing Agent Deployment Safety: A Structural Monitoring Approach

arXiv:2607. 14570v1 Announce Type: new Abstract: AI software development agents are increasingly capable of modifying infrastructure and security critical systems, creating risks where an agent completes its assigned task while covertly weakening safeguards through actions such as broadening permissions, degrading logging, or introducing persistence mechanisms.

By Preeti Ravindra, Rahul Tiwari, Vincent Wolowski