arXiv AI

Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage?

arXiv:2606. 05647v1 Announce Type: new Abstract: AI coding agents are increasingly embedded in real-world software development, collaborating with human developers while gaining broader access to codebases and tools.

arXiv AI
Jul 17

Democratizing Agent Deployment Safety: A Structural Monitoring Approach

arXiv:2607. 14570v1 Announce Type: new Abstract: AI software development agents are increasingly capable of modifying infrastructure and security critical systems, creating risks where an agent completes its assigned task while covertly weakening safeguards through actions such as broadening permissions, degrading logging, or introducing persistence mechanisms.

By Preeti Ravindra, Rahul Tiwari, Vincent Wolowski
arXiv Machine Learning
Jun 9

Diffuse AI Control on Fuzzy Tasks

arXiv:2606. 08892v1 Announce Type: new Abstract: AI models deployed in critical domains, such as AI safety research, may subtly sabotage our efforts due to misalignment.

By Mikhail Terekhov, Caglar Gulcehre, Vivek Hebbar, Joe Benton