Watchtower

Watchtower tracks when AI companies change their safety policies — and when they break them.

Anthropic

Published a secondary compliance framework, substituting the RSP for the sake of new compliance requirements from the EU and California

Major
Change

Anthropic

Dec 19, 2025

Anthropic published a separate Frontier Compliance Framework (FCF) to satisfy California’s SB 53, which takes effect January 1, 2026. Rather than using its Responsible Scaling Policy (RSP), which is built on concrete if-then commitments that tie specific safeguards to capability thresholds, Anthropic created a thinner document, stripped of those hard commitments. The FCF defines risk tiers for cyber offense, CBRN, and loss-of-control scenarios, but attaches no binding mitigations to them, stating instead that “the specific mitigations we implement may be determined when the relevant risk tier is reached.” The FCF does not include much of what gave the RSP its rigor: specific mitigations tied to risk thresholds, detailed Capability and Safeguard Reports, delegation of decision-making authority on model releases to the Responsible Scaling Officer, and the involvement of Anthropic’s Long-Term Benefit Trust.

The only safety framework that regulators can enforce is the one with the fewest commitments. When Anthropic endorsed SB 53, it wrote, “These requirements would formalize practices that Anthropic and many other frontier AI companies already follow. At Anthropic, we publish our Responsible Scaling Policy, detailing how we evaluate and mitigate risks as our models become more capable … Now all covered models will be legally held to this standard.”

If one of the most safety-conscious AI companies responds to legislation it supported (and characterized as a codification of its existing practices) by publishing a separate, thinner safety framework, essentially escaping the codification of its existing practices, other companies will take note and may do the same thing.

OpenAI

Sep 9, 2026

Updated the GPT-6 Astra system card, specifically to hedge claims about misalignment.

View details

OpenAI

Aug 18, 2026

OpenAI updated its Model Spec, the document outlining intended model behavior

View details

xAI

Aug 17, 2026

xAI updated Grok 4.6's model card after release. A changelog was included.

View details

Google

Aug 14, 2026

Google updated its Gemini 3.7 Flash model card after publication, making changes to language in the “Key Results for Gemini 3.7 Flash” column in the Frontier Safety Assessment section

View details

OpenAI

Aug 3, 2026

OpenAI updated its system card for two of its models: GPT 5.6 and GPT Live.

View details

Anthropic

Jul 24, 2026

Anthropic updated its Frontier Compliance Framework.

View details

Frequently asked questions

Have more questions? Our team is happy to help, contact us.

What does The Midas Project do?
What does your name mean?
Who is behind The Midas Project?
Are you anti-AI?
How can I contribute?
How can I get in touch?