Self-regulation by leading AI companies is insufficient to manage the safety risks of increasingly capable systems
What's this about?
People disagree about whether AI firms can keep their own powerful AI systems safe. The key question is whether their safety steps work without outside checks and firm laws.
What supporters say
- AI firms control much of the proof used to judge if their systems are safe.
- Firm laws can add clear duties, outside checks, and punishments when firms break rules.
- Company promises often lack strong checks, clear fines, and real power to make firms obey.
- More skilled AI can act in ways that safety tests may miss or fail to spot.
What critics say
- Voluntary safety rules can change faster than laws, which may take years to pass.
- In-house safety plans can set clear steps for testing, fixing, and watching AI systems.
How to read this
The number of points on each side does not show who is right; the strength of each point matters more.
The bottom line
Company safety steps help, but they cannot manage all the risks on their own. The evidence leans toward using outside checks and firm laws too.
The claim is that voluntary safety measures by leading AI companies cannot, on their own, reliably control the risks posed by increasingly capable systems. The evidence points toward self-regulation as a useful layer of protection, but not a complete substitute for independent oversight and binding rules.
The case for
The strongest argument concerns accountability. Voluntary promises usually lack dependable outside enforcement, meaningful penalties, and independent checks. An analysis of companies’ commitments to the U.S. government found uneven implementation and limited transparency. Frontier AI commitments also leave much of the interpretation, implementation and enforcement to the companies themselves. 1
Companies also control much of the evidence used to support their own safety claims. They often design, carry out and disclose the evaluations that assess their systems. Research has found inconsistent methods, limited comparability and incomplete reporting of results before and after safety measures are applied. The Foundation Model Transparency Index likewise found uneven disclosure about systems’ capabilities, limits, risks and effects on users and society (see Figure 1). 2
There is a further problem: advanced AI systems are difficult to evaluate with confidence. The International AI Safety Report says current tests are imperfect and that safeguards can fail or be bypassed. The AI Index has recorded rising numbers of reported incidents and controversies, while standardized reporting remains limited (see Figure 2). 3 This makes it harder for outsiders to know whether a company’s controls are working as promised.
Binding regulation cannot eliminate these risks, but it can add duties and consequences that voluntary programs lack. The European Union’s AI Act, for example, requires specified systems to meet obligations involving documentation, risk management, evaluation and oversight. Such laws reflect a move from voluntary assurances toward enforceable responsibilities. 4
The case against
Self-regulation can still produce specific and technically informed safety controls, and it may move faster than legislation. Anthropic, OpenAI and Google DeepMind have published frameworks that link capability tests or risk thresholds to safeguards, deployment limits, mitigation plans and escalation procedures.
The voluntary NIST framework provides a common structure for identifying, measuring and managing AI risks. International commitments can also establish shared expectations before detailed laws are in place. These efforts show that companies can create operational safety processes, rather than merely issue broad promises. 5
Voluntary standards may also have a practical advantage: they can be updated as technology changes, while legislation often takes years to pass and implement. 6 But the evidence does not show that these internal systems are consistently applied, independently checked or sufficient for risks that extend beyond a company’s own operations.
The bottom line
The evidence favours the claim, but only with low confidence. The better-supported findings concern missing enforcement, limited transparency and the weaknesses of current evaluations. The opposing evidence shows that internal controls exist and may be useful, but it does not establish that self-regulation alone reliably manages the risks of increasingly capable AI.
The most defensible conclusion is conditional. Self-regulation should remain one part of risk management, but it needs to be supplemented by independent oversight and binding rules. Provider-authored frameworks differ in their level of detail, risk thresholds, testing methods and governance arrangements, while differences in safety reporting make results difficult to compare (see Figure 3).
Important uncertainties remain. The available evidence describes policies, reporting gaps and governance structures more clearly than it measures whether particular controls reduce real-world harm. Reported incidents cannot by themselves show which governance system caused or prevented a failure. The central unresolved issue is whether outsiders can verify that companies apply their policies effectively—and whether firms face enough consequences when they do not.
Figures & data

All contributions are reviewed for clarity, balance, and evidence. The strongest insights are elevated into the argument graph — with credit to you.
Help improve this analysis →