120 Organisations Are Building AI Cybersecurity Rules. Here Is What SAFE Means.
Key takeaways
- The Open Secure AI Alliance has grown to more than 120 member organisations and is developing SAFE cybersecurity transparency guidelines for agentic AI
- Agentic AI systems that can execute code, browse the web, and connect to APIs present novel cybersecurity risks that current frameworks do not adequately address
- The guidelines focus on how AI systems disclose and handle security-relevant situations, but details of testable requirements are not yet public
The Open Secure AI Alliance has grown quickly. What started as a coalition focused on defining honest and trustworthy agentic AI has now reached more than 120 member organisations, and the group's latest work is specifically targeting cybersecurity transparency. The new guidelines being developed carry the name SAFE, and they are aimed at strengthening how agentic AI systems handle and disclose cybersecurity-relevant information.
The timing is not accidental. Agentic AI, systems that take actions autonomously on behalf of users or organisations, are being deployed faster than security frameworks have been developed to govern them. An agent that can browse the web, execute code, connect to APIs, and manage files on behalf of a user is also an agent that can be compromised, manipulated, or used as a vector for attacks. The question of how these systems should behave when they encounter security-relevant situations, and how they should disclose that behaviour, is genuinely unresolved.
What Transparency Means in Cybersecurity AI
Cybersecurity transparency for AI systems covers a range of concerns. At the most basic level, it includes questions like: does the system tell users when it has encountered a potential security threat? Does it disclose what actions it has taken in response? Does it behave consistently regardless of whether it thinks it is being observed or tested?
At a deeper level, transparency concerns include how AI systems handle vulnerabilities they discover, whether they report security issues through appropriate channels, and how they respond to adversarial inputs designed to manipulate their behaviour. These are not hypothetical edge cases. As AI agents become more capable and more widely deployed, they will increasingly be the first line of detection for security incidents, and their behaviour in those moments will matter enormously.
The SAFE acronym has not been fully detailed in the available information, but the direction is clear. The alliance is trying to establish shared standards that AI developers and deployers can use to signal that their systems meet minimum requirements for cybersecurity transparency. The goal is presumably to avoid a situation where every organisation deploying agentic AI invents its own definitions and claims of safety that cannot be compared or evaluated externally.
Why an Alliance Model Makes Sense Here
Cybersecurity standards have historically been difficult to establish precisely because organisations often resist transparency that could expose their vulnerabilities. The alliance model, where more than 120 organisations collectively develop and presumably commit to shared standards, has a better chance of creating genuine adoption than a top-down regulatory approach.
The 120-plus membership is also interesting as a signal. When that many organisations are willing to participate in developing shared standards, it suggests there is meaningful consensus that the problem is real and that voluntary action is preferable to waiting for regulation. In cybersecurity, that kind of proactive industry consensus is genuinely unusual. The field has more often operated on the principle of disclosing as little as possible.
The involvement of NVIDIA in this effort is worth noting, since the SAFE guidelines are being highlighted on the NVIDIA Newsroom. NVIDIA's position at the infrastructure layer of AI deployment gives it a particular interest in security standards that apply to the systems running on its hardware. If agentic AI systems built on NVIDIA infrastructure are involved in security incidents because of unclear standards, that reflects on the platform.
What Happens Next
The guidelines are still being developed, which means the detail that would allow real evaluation is not yet public. The critical questions are whether the SAFE standards will include testable, verifiable requirements or remain at the level of principles, and whether there will be any mechanism for independent audit or certification.
Principles-based guidelines in tech have a mixed track record. They can establish useful common vocabulary and raise the floor across an industry, but without enforcement or verification they can also become a form of credentialled compliance theatre. The alliance has the membership to make something meaningful. Whether the SAFE guidelines end up with teeth will determine how significant this effort actually is.