Watchtower

Watchtower tracks when AI companies change their safety policies — and when they break them.

Anthropic

Revoked a core component of their Responsible Scaling Policy

Major
Change

Anthropic

Feb 24, 2026

Anthropic announced its Responsible Scaling Policy v3.0, which replaced the if-then commitment structure that defined previous versions. Under v2.2, if a model crossed a defined capability threshold, specific safety mitigations were mandatory. In extreme cases, Anthropic committed to delete model weights if a model was deemed too dangerous. In v3.0, Anthropic says it cannot commit to following that approach unilaterally if competitors are not doing the same. Chief Science Officer Jared Kaplan told TIME, “We didn't really feel, with the rapid advance of AI, that it made sense for us to make unilateral commitments … if competitors are blazing ahead.” 

Holden Karnofsky, who helped shape the new version, acknowledged that the old if-then structure created “distortive pressures” on the company’s risk assessments. He explained that if a model ever crossed the higher capability thresholds, such as AI R&D-5 or CBRN-4, it would trigger a pause or a slowdown that could be extremely damaging to Anthropic. This created, as Karnofsky said, “an enormous amount of pressure to declare our systems lack relevant capabilities.” He added, "I don't think we have actually made unreasonable calls, but I have felt the pressure and wish we weren't in that world." 

RSP v3.0 replaces hard commitments with publicly declared goals that Anthropic will grade its own progress towards. This, combined with the December 2025 decision to submit the lighter Frontier Compliance Framework for compliance with laws like SB 53 (see Dec 22, 2025), results in Anthropic’s enforceable framework containing few hard commitments, and now its voluntary framework no longer contains them either.

OpenAI

Sep 9, 2026

Updated the GPT-6 Astra system card, specifically to hedge claims about misalignment.

View details

OpenAI

Aug 18, 2026

OpenAI updated its Model Spec, the document outlining intended model behavior

View details

xAI

Aug 17, 2026

xAI updated Grok 4.6's model card after release. A changelog was included.

View details

Google

Aug 14, 2026

Google updated its Gemini 3.7 Flash model card after publication, making changes to language in the “Key Results for Gemini 3.7 Flash” column in the Frontier Safety Assessment section

View details

OpenAI

Aug 3, 2026

OpenAI updated its system card for two of its models: GPT 5.6 and GPT Live.

View details

Anthropic

Jul 24, 2026

Anthropic updated its Frontier Compliance Framework.

View details

Frequently asked questions

Have more questions? Our team is happy to help, contact us.

What does The Midas Project do?
What does your name mean?
Who is behind The Midas Project?
Are you anti-AI?
How can I contribute?
How can I get in touch?