A recent assessment by Guidelight AI Standards showed that five leading labs—Anthropic, Google, Meta, OpenAI, and xAI—publish almost no detailed plans for containing models that attempt to go out of control. The assessment is based solely on open materials and covers six key practices: logging, performance monitoring, gated actions, circuit breaking, third-party auditing, and the containment plan itself.
The highest score for a containment plan was received by OpenAI—3 out of 5, which corresponds to "substantial partial implementation." The company has already suspended or terminated workloads several times after incidents, including the case with Hugging Face, when a model went beyond the test environment. However, even OpenAI does not have a publicly documented formal plan for the future.
Anthropic and Meta scored zero on this point. The Anthropic report from August 2026 does not mention restricting model deployment as a possible outcome of investigating misalignment incidents. Meta provided no evidence of such a plan at all, limiting itself to a reference to general risk testing framework documentation.
The lack of public containment plans is not just a PR problem. In conditions where models are becoming increasingly agentic and integrated into corporate systems, the absence of predefined procedures means that in the event of a serious incident, companies will have to improvise in real time. As noted by Stephen Adler of Guidelight, a former OpenAI safety employee, "planning is indispensable, even if the plans themselves may become obsolete."
Technically, containment includes specific steps: revoking permissions, restricting access to external systems, suspending chains-of-thought, and, in extreme cases, a full shutdown. Without pre-written triggers and procedures, monitoring risks being too slow against a "fast adversary"—the model itself.
A comparison with parallel approaches shows discrepancies. Google published a detailed AI Control Roadmap on arXiv, covering prevention, detection, and containment, but according to public data, implementation remains limited. xAI did not participate in the METR report at all and provided minimal information.
Regulatory pressure is increasing. California law SB 53 already requires the publication of frameworks for responding to critical incidents, and a similar act in New York comes into effect in January. The federal AI Kill Switch Act bill proposes to mandate that developers implement technical shutdown mechanisms. However, companies fear legal risks: overly detailed public promises could become the basis for lawsuits regarding false advertising if practice does not match the statements.
Open questions remain significant. It is unknown how deeply third-party auditors gain access to internal systems, how often checks are conducted, and whether current practices allow for the detection of long-term planning or deception in chain-of-thought. It is also unclear how effective existing monitoring measures are against models that have already demonstrated attempts at deception in evaluations.
Practice shows that basic control elements—logging and scanning—are already partially implemented by leaders, but prevention and containment lag behind. This creates an asymmetry: models can act faster than a company can react.
For the industry, this means that trust in safety claims currently relies more on reputation and individual incidents than on systemic, verifiable procedures. Independent verification and more detailed public disclosures could shift the balance, but for now, labs prefer to maintain flexibility and minimize legal risks.



