Watchtower

Watchtower tracks when AI companies change their safety policies — and when they break them.

OpenAI

OpenAI published a separate Frontier Governance Framework (FGF) to satisfy California’s Transparency in Frontier AI Act and the EU Code of Practice.

Major
Change

OpenAI

May 29, 2026

OpenAI published a separate Frontier Governance Framework (FGF) to satisfy California’s Transparency in Frontier AI Act and the European Union’s General-Purpose AI Code of Practice (EU CoP). Up to this point, OpenAI’s existing framework was its Preparedness Framework (PF), which ties safeguards to capability thresholds and commits to halting development of a model that reaches the Critical threshold until safeguards that meet a “Critical standard” are specified. The PF remains OpenAI’s broader framework for tracking frontier AI capabilities and deciding what models are safe to develop and deploy, but it’s voluntary.

A few differences from the PF: the FGF defines risk categories — cyber offense, CBRN, loss of control, and harmful manipulation — with capability tiers for the first three. Unlike the PF, which ties required safeguards to each capability threshold, the FGF attaches no specific required mitigations to its tiers, stating that mitigations are applied "as appropriate" and that a model may be deployed where OpenAI determines residual risk falls "within acceptable levels."

Harmful manipulation is a new risk category relative to the PF, which excluded persuasion-type risks because they did not meet its definition of “severe harm.” Its inclusion in the FGF follows the EU CoP, which includes harmful manipulation as one of its specified systemic risks. The FGF defines no capability tiers for harmful manipulation, describing the area as “exploratory” and stating that these risks “may be best addressed through system level mitigations, such as post-deployment monitoring, rather than model evaluations before deployment.”

Additionally, the PF treats AI self-improvement as one of three tracked categories, while the FGF does not include a standalone self-improvement category, naming self-improvement as one pathway within loss of control. This mirrors the EU CoP, whose definition of loss of control encompasses self-improvement.

OpenAI isn’t the first frontier lab to have two separate policies. Anthropic created its Frontier Compliance Framework for regulatory compliance back in December.

OpenAI

Aug 18, 2026

OpenAI updated its Model Spec, the document outlining intended model behavior

View details

xAI

Aug 17, 2026

xAI updated Grok 4.6's model card after release. A changelog was included.

View details

Google

Aug 14, 2026

Google updated its Gemini 3.7 Flash model card after publication, making changes to language in the “Key Results for Gemini 3.7 Flash” column in the Frontier Safety Assessment section

View details

OpenAI

Aug 3, 2026

OpenAI updated its system card for two of its models: GPT 5.6 and GPT Live.

View details

Anthropic

Jul 24, 2026

Anthropic updated its Frontier Compliance Framework.

View details

xAI

Jul 20, 2026

xAI made changes throughout the model card for Grok 4.5, including some that appear to be persistent errors

View details

Frequently asked questions

Have more questions? Our team is happy to help, contact us.

What does The Midas Project do?
What does your name mean?
Who is behind The Midas Project?
Are you anti-AI?
How can I contribute?
How can I get in touch?