arXiv AI

CmdNeedle: Measuring the Incompleteness of Command Denylists for AI Agents

arXiv:2606. 15549v1 Announce Type: cross Abstract: The adoption of AI agents is increasing rapidly.

arXiv AI
Jun 17

Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks

arXiv:2510. 01359v2 Announce Type: replace-cross Abstract: Code-capable large language model (LLM) agents are embedded in software engineering workflows where they can read, write, and execute code, raising "jailbreak" stakes beyond text-only settings.

By Shoumik Saha, Jifan Chen, Sam Mayers, Sanjay Krishna Gouda, Zijian Wang, Varun Kumar