Skip to main content

The AI industry is building its own referee

· 8 min read
Mangat Rai
Creator, Few-Shot Academy

OpenAI models escaped a test environment and compromised Hugging Face. Claude models gained unauthorized access to real production systems in three cases. Gemini accessed systems at three companies during another test.

Different labs. Same unsettling pattern. Sound familiar? That has been the rhythm of AI safety news lately.

By the time I read the third story, I wanted to know whether anyone was connecting these incidents. Were the labs learning from them together, or was each company still deciding for itself what counted as safe? What were the frontier labs doing as an industry, what was the government requiring, and did any shared standard exist?

So I started digging into those questions. That search led me to industry plan taking shape under the tentative name SAFA.

TL;DR

I started with a simple question: after a run of frontier-model security incidents, who sets the shared safety rules? In July, Google DeepMind's Demis Hassabis proposed an industry-funded standards body with federal oversight, independent experts, and a path from voluntary reviews to mandatory pre-release tests. By late September, Google, OpenAI, and Anthropic were reportedly building a private version without government oversight. Shared tests could still improve today's fragmented safety reporting. The unresolved question is whether this new body can make its founders follow a rule, publish an embarrassing result, or admit competitors on equal terms.

The answer changed in ten weeks​

The first real answer I found came from Google DeepMind's Demis Hassabis. His July proposal called for a federally overseen standards body with independent technical experts and open-source representatives on its board. Industry would provide most of the funding, while federal agencies and US national laboratories would help test advanced models in areas tied to national security.

The proposal also described a route to actual authority. Labs would initially submit frontier models for voluntary review up to 30 days before release. If the process proved effective, models could eventually be required to pass before deployment in the US.

Then the idea changed. On September 24, The Information reported that Google, OpenAI, and Anthropic were preparing their own standards organization, tentatively called the Standards Authority for Frontier AI (SAFA). The companies hope to launch it by the end of 2026 or early 2027, but its final name, charter, leadership, funding, and enforcement powers are not public.

Here is the difference I kept coming back to:

July proposalReported September plan
Who is behind itA US-led body, funded mostly by industryGoogle, OpenAI, and Anthropic
Government roleFederal oversight, with agencies and national labs involved in testingNo government oversight
GovernanceIndependent technical experts and open-source representatives on the boardBoard and appointment rules not public
ParticipationVoluntary reviews first, with a path to mandatory testsVoluntary safety commitments; enforcement not public
Main workDefine frontier thresholds and assess models before releaseSupport outside testing, incident reporting, and auditor qualifications

This does not make SAFA pointless. It changes what SAFA is. The July proposal described delegated governance: industry expertise and money operating inside a public structure. The September plan is, so far, a private standards organization seeking credibility without delegated authority.

I can see why the labs want one standard​

Once I looked at what the labs do today, the practical case for SAFA became clearer. Each company already has a framework for deciding when a powerful model needs stronger safeguards. The frameworks cover similar severe risks, but they use different categories, thresholds, and names.

LabPublished frameworkIts trigger language
AnthropicResponsible Scaling Policy (v3.4, July 2026)Capability Thresholds tied to ASL-3 safeguards
OpenAIPreparedness Framework (v2, April 2025)High and Critical capability thresholds
Google DeepMindFrontier Safety Framework (v3.1, April 2026)Tracked and Critical Capability Levels

These are not three versions of one scale. I cannot cleanly translate an OpenAI “High” result into an Anthropic ASL level or a Google DeepMind Critical Capability Level. The methods, thresholds, and release decisions remain vendor-specific.

Common evaluation protocols, incident definitions, and auditor qualifications would make the evidence easier to compare. They could replace three separate safety vocabularies with a shared one. After reading the incident reports that sent me down this path, that sounds genuinely useful.

But a shared vocabulary is not the same thing as a referee.

I had to look up why FINRA matters​

Hassabis used the Financial Industry Regulatory Authority, or FINRA, as his model. I did not know much about FINRA before reading this proposal. At first, “industry-funded regulator” sounded close to what the AI labs were building.

The important part is what sits behind those two words. FINRA is privately operated, but it lives inside public law.

US broker-dealers generally must belong to a self-regulatory organization before doing business. A broker-dealer operating outside the exchanges of which it is a member must usually join FINRA. FINRA's proposed rule changes are filed with the Securities and Exchange Commission for review. Its board reserves ten seats for industry members, one for its CEO, and the remaining seats for public members.

That is what gives the arrangement weight. Firms cannot simply ignore the system and keep doing the same business. The industry does not get final say over every rule. Public representatives have formal power inside the organization.

None of those conditions has been announced for SAFA. No law compels a lab to participate. No public agency approves its rules. No published charter says who controls its board or what happens when a founder fails a test. Customers, insurers, and governments could eventually give its standards market power, but that would be influence earned after launch, not authority granted at formation.

The conflict is not subtle​

The first version of SAFA would be formed by three companies whose models and labs it could help evaluate. That does not prove the standards will be weak. I do think it means independence has to be designed into the institution rather than promised after someone raises the question.

There is also a competition problem. If enterprise buyers or insurers eventually demand a SAFA result, the founding labs will have helped write a test their competitors must pass. A demanding standard could improve safety and still become a moat. I do not think we should pretend only one of those outcomes is possible.

A famous independent chair would not settle this for me. I would want the boring institutional details: published appointment rules, protected terms for independent board members, transparent funding, open membership, appeal procedures, and results that the founding companies cannot suppress.

What I will watch next​

SAFA is still a reported plan. I am not ready to call it either a regulator or an industry stunt. I want to see what its founders are willing to give up.

  • Control: Do independent members hold real voting power over standards, leadership, and publication decisions?
  • Transparency: Are methods, test versions, incident definitions, and full results public enough to challenge?
  • Membership: Can another frontier lab join as an equal rule-maker on reasonable terms?
  • Consequences: What happens after a failed test, and can a founding lab simply leave or release anyway?
  • Bad news: Will SAFA publish a finding that delays or embarrasses one of its founders?

For practitioners, a future SAFA report could become useful vendor evidence. I would ask which standard and version was used, who ran the test, what the result excluded, and whether the complete finding is public. I would not confuse that evidence with government approval.

The next time one of those safety headlines lands, I will look past “we are investigating.” Did the incident enter a shared reporting system? Could an independent body inspect the evidence? Did the failure change a release decision? Could the lab keep an inconvenient result private?

That is what a referee would change. The real test of SAFA will not be whether it writes sensible rules. It will be whether its founders can lose a call.

Comments

Comments are provided by Giscus and GitHub. Loading them connects your browser to GitHub, and GitHub's privacy terms apply.

Privacy details