SECURITY

The Open Secure AI Alliance Wants Every AI System to Have a Safety Label

(1 month ago) · 5 min read

Key takeaways

  • The Open Secure AI Alliance has more than 120 member organisations developing the SAFE cybersecurity transparency framework
  • The framework is specifically oriented toward agentic AI systems, which present a different attack surface from passive models
  • Prompt injection attacks, where malicious content hijacks an agent's actions, are among the key threat vectors the framework aims to address
  • Currently there is no standard security disclosure format that AI vendors are expected to follow for enterprise procurement

More than 120 organisations are now part of the Open Secure AI Alliance, and they are working on something that could become the closest thing the AI industry has to a standardised safety transparency framework. The initiative, called SAFE (Secure AI Framework for Enterprises, or similar, with the full acronym pending formal publication), aims to give organisations a consistent way to assess and communicate the cybersecurity properties of AI systems, particularly agentic AI that takes autonomous actions.

The announcement came via NVIDIA's newsroom in early August 2026, and while the alliance already has significant membership depth spanning AI developers, cloud providers, enterprise software companies, and cybersecurity firms, the guidelines themselves are still in development. What is known is that the framework is oriented specifically toward agentic AI systems, which present a distinct security profile compared to passive AI models.

Why Agentic AI Needs Its Own Security Framework

Most existing cybersecurity frameworks were built around systems that respond to human requests rather than taking autonomous actions. A traditional enterprise application does something when a human tells it to. An agentic AI system can browse the web, call APIs, write and execute code, send emails, and interact with external services based on its own planning and decision-making. The attack surface is fundamentally different.

Consider the category of prompt injection attacks, which have become one of the most concerning threat vectors for agentic AI. A malicious actor embeds instructions in content that an AI agent is likely to process, such as a webpage it browses, a document it reads, or an email it handles. The agent, treating that content as legitimate input, follows the embedded instructions rather than the user's actual intent. The result can range from data exfiltration to the agent taking destructive actions in external systems it has access to.

Existing security standards do not provide clear guidance on how to evaluate whether an AI system is robust against this kind of attack. The SAFE guidelines appear to be aiming at exactly this gap, defining what transparency organisations should expect from AI vendors about their systems' security properties, and what testing and disclosure standards should look like.

What Standardisation Would Actually Change

Right now, if an enterprise wants to deploy an AI agent with access to its internal systems, it has to conduct its own security evaluation largely from scratch. There is no standard questionnaire, no common set of benchmarks, no agreed disclosure format that AI vendors are expected to follow. Every procurement is a bespoke security assessment, which is expensive, slow, and inconsistent.

The future, in 3 minutes a day. The biggest tech story explained every morning, free. Get the briefing →

A widely adopted SAFE framework would change that. If AI vendors were expected to publish standardised security disclosures covering known vulnerabilities, testing methodologies, update and patching practices, and incident response procedures, enterprise security teams would have a common baseline to work from. That reduces evaluation costs and, more importantly, makes it harder for vendors to obscure security shortcomings behind opaque marketing language.

For the AI vendors themselves, there is a credibility argument for participating. Companies that can demonstrate compliance with a recognised framework have a competitive advantage in enterprise sales cycles where security is a gating factor. The 120-plus member count suggests the alliance has enough breadth to potentially make this a de facto standard even without formal regulatory backing.

The Limitations to Watch For

Frameworks developed by industry alliances have a predictable failure mode: they get designed around what members are comfortable disclosing rather than what would genuinely inform security decisions. The history of voluntary industry standards is littered with examples of frameworks that create the appearance of accountability without the substance.

The involvement of AI developers alongside enterprise customers and cybersecurity specialists is a mitigating factor here, but the real test will come when specific disclosure requirements are published. If vendors are expected to share detailed information about known failure modes, red-team results, and adversarial robustness testing, that is a meaningful commitment. If the framework reduces to a self-attestation checklist, it will provide limited real-world protection.

The alliance also needs to think about how the framework handles AI systems that are updated frequently. A security disclosure made for a model at one version may not reflect the properties of the next release. Dynamic systems require dynamic transparency, which is harder to standardise than a one-time disclosure document. That challenge will likely define how useful SAFE turns out to be in practice.

More from Future Technology