Open Secure AI Alliance Proposes SAFE Guidelines for Agentic AI Cybersecurity
Key takeaways
- The Open Secure AI Alliance has grown to more than 120 member organisations across the tech industry
- The SAFE guidelines focus on cybersecurity transparency for agentic AI systems, which can autonomously take real-world actions
- Prompt injection attacks, where malicious instructions are embedded in content an AI agent processes, have been demonstrated against commercial agentic systems in 2025 and 2026
- The framework is voluntary but may be referenced by EU, UK, or US regulators, giving it effective regulatory force
As AI agents become more capable of taking real actions in the world, from booking appointments to executing code to managing infrastructure, the question of how to secure them has become genuinely urgent. The Open Secure AI Alliance, now comprising more than 120 organisations, is stepping into that gap with a new framework: the SAFE guidelines for cybersecurity transparency in agentic AI systems.
The timing is not coincidental. Agentic AI, meaning AI systems that can autonomously plan and execute multi-step tasks with real-world consequences, has moved from research curiosity to commercial deployment remarkably quickly. Major software vendors are shipping agentic features into enterprise tools, and the security community has been raising alarms about a class of attacks that traditional cybersecurity defences were simply not designed to handle.
What Makes Agentic AI a Distinct Security Problem
Conventional software security is built around a relatively clear model: protect data, control access, monitor for known attack patterns. Agentic AI breaks several assumptions that model relies on. An AI agent acting on behalf of a user might have access to email, documents, code repositories, and external APIs simultaneously. It makes decisions dynamically, often without a human reviewing each step. And it can be manipulated through the data it processes, not just through direct system access.
Prompt injection, for instance, is an attack where malicious instructions are embedded in content that an AI agent reads and then follows, potentially overriding the legitimate instructions it was given. This is not a theoretical concern. Security researchers have demonstrated prompt injection attacks against several commercial agentic systems in 2025 and 2026. When an AI agent has the authority to send emails, make purchases, or modify files, a successful prompt injection can have immediate, material consequences.
The SAFE framework, as outlined by the Alliance, focuses on cybersecurity transparency, which suggests its primary goal is not to prescribe specific technical controls but to establish what information AI developers, deployers, and users should disclose about how their agentic systems handle security risks. Transparency frameworks of this kind are typically a precursor to more binding standards, giving the industry time to converge on common practices before regulators move in.
Why 120 Organisations Matters
The breadth of the Alliance is worth noting. With more than 120 member organisations, this is not a small working group of like-minded vendors writing guidelines that only benefit themselves. A coalition of this size typically includes large enterprise software companies, cloud providers, cybersecurity firms, and research institutions, creating enough diversity of interest that the resulting guidelines carry some credibility.
NVIDIA's involvement, given that its platforms underpin much of the agentic AI being deployed commercially, gives the Alliance particular weight on the infrastructure side. The question that always arises with voluntary industry guidelines is enforcement. A transparency framework that companies can self-certify against has limited teeth. The more interesting signal will be whether governments in the EU, UK, or US choose to reference the SAFE guidelines in their own AI regulatory frameworks, which would effectively give them regulatory force.
The Bigger Picture
What the SAFE guidelines represent is the industry acknowledging, in a coordinated and public way, that agentic AI creates security risks that are new and not yet fully understood. That acknowledgement alone is significant. For much of the past two years, the dominant narrative from AI developers has been focused on capability and productivity gains, with security concerns treated as edge cases or future problems.
The fact that more than 120 organisations are now collaborating on a formal security transparency framework suggests the industry recognises that the edge cases are arriving faster than expected. For enterprises currently evaluating or deploying agentic AI, the SAFE guidelines will be worth watching closely. Whether or not your vendor is a member of the Alliance, the framework is likely to shape what questions procurement and security teams should be asking.