FLARE-AI: Flaw Reporting for AI
arXiv:2606. 31567v1 Announce Type: cross Abstract: Flaw reporting for deployed AI systems is fundamental to identifying system failures and improving AI safety.
The paper discusses the need to adapt incident reporting frameworks for AI agents, which are rapidly deployed and face unique security challenges. By comparing AI systems and agents and consulting 23 experts, the authors identify key reporting elements such as agent memory, autonomy levels, and tool usage. They also highlight open research questions, potential reporting weaknesses like data leakage, and outline privacy requirements for secure AI agent deployment.
arXiv:2606. 31567v1 Announce Type: cross Abstract: Flaw reporting for deployed AI systems is fundamental to identifying system failures and improving AI safety.
arXiv:2606. 25836v2 Announce Type: replace Abstract: To better assist users with completing challenging tasks, AI agents mediate communications, access data, and interact with different APIs.
To better assist users with completing challenging tasks, AI agents mediate communications, access data, and interact with different APIs. Many employers (and even nation-states) already provide their users with this technology.
arXiv:2607. 25379v1 Announce Type: new Abstract: Cyber-capable AI agents combine language models with tools, memory, and execution en- vironments to perform multi-step offensive-security tasks.
arXiv:2606. 17114v1 Announce Type: cross Abstract: AI agents are increasingly being adopted in enterprise and personal settings with access to emails, databases, documents, and other tools where they can read, update, and disseminate sensitive information.
arXiv:2607. 05163v1 Announce Type: cross Abstract: AI systems may produce failures after deployment that pre-deployment safety assessments do not anticipate.
AI agents increasingly act through tools and delegated authority, but general incident repositories rarely capture the mechanisms needed to compare public failures with agent-security evaluations. We...
The paper discusses the trustworthiness of agentic AI systems built on large language models, highlighting new security and operational risks such as indirect prompt injection, memory contamination, and cross‑session data leakage. It categorizes failure modes, reviews mitigation strategies—including instruction hierarchies, context isolation, and constrained tool use—and introduces the Trustworthy Agent Development Lifecycle (TADL), a six‑phase framework for specification, design, training, evaluation, deployment, and monitoring. The authors note that TADL has not yet been empirically validated but offers a structured foundation for developing more secure and accountable agentic systems, and they call for improved benchmarks and future research priorities.
The paper proposes AI Deployment Accountability Engineering (ADAE), a new subdiscipline focused on establishing measurable, continuous, and actionable accountability for AI systems once they are deployed. ADAE treats accountability as a deployment-layer property, aiming to ensure systems remain within acceptable risk limits, identify failure contexts, attribute failures across technical and human components, and translate technical failures into downstream consequences. The authors outline a research agenda built around four pillars—structured discovery of context-dependent failure modes, privacy-preserving accountability measurement, system-level risk analysis for agentic AI, and translation of technical failures into operational and institutional risks—to support timely intervention in safety-critical socio-technical environments.
The article proposes a structured framework of behavioral indicators that could signal a progression toward potentially catastrophic threats from AI systems. Drawing on established methods from cybersecurity and national security, it defines clear metrics, indicators, and thresholds across multiple dimensions of AI capability and behavior. The framework is intended to enable researchers and policymakers to implement evidence‑based monitoring protocols for rogue AI progression.
arXiv:2609. 11030v1 Announce Type: new Abstract: AI agents increasingly act through tools and delegated authority, but general incident repositories rarely capture the mechanisms needed to compare public failures with agent-security evaluations.
arXiv:2609.23894v1 Announce Type: cross Abstract: Agentic AI extends LLM security beyond generated content to persistent state, autonomous actions, tool use, and interactions with humans and other ag...