Watchtower

Watchtower tracks when AI companies change their safety policies — and when they break them.

Anthropic

Anthropic removed its commitment to define ASL-4 evaluations before reaching ASL-3, and released an ASL-3 model without some promised safeguards.

Moderate
Violation

Anthropic

May 14, 2025

Anthropic's original 2023 RSP committed to "define ASL-4 evaluations before we first train ASL-3 models," including both capability thresholds and "warning sign evaluations." Anthropic removed this commitment in RSP version 2.1, which was noted in the PDF but not the changes they shared on the web.

When Claude Opus 4 was deployed under ASL-3, Anthropic stated that capability thresholds for ASL-4 are now defined. However, the RSP still lacks the warning sign evaluations originally promised. Additionally, the AI R&D-5 threshold triggers ASL-4 security but says nothing about ASL-4 deployment standards—a notable gap given that models with advanced AI R&D capabilities would be highly tempting to deploy internally, and internal deployments are in scope for deployment standards under Anthropic's RSP.

Nine days before announcing Claude Opus 4 as an ASL-3 model, Anthropic revised their RSP to lower ASL-3 security requirements. The change removed the requirement to be robust against employees with access to "systems that process model weights" attempting to steal model weights—potentially a large fraction of technical staff. Anthropic justified the change by arguing model theft isn't central to CBRN-3 or AI R&D-4 risks. Critics argue the timing suggests potential prior noncompliance, and that security remains critical as models approach ASL-4 thresholds.

OpenAI

Aug 18, 2026

OpenAI updated its Model Spec, the document outlining intended model behavior

View details

xAI

Aug 17, 2026

xAI updated Grok 4.6's model card after release. A changelog was included.

View details

Google

Aug 14, 2026

Google updated its Gemini 3.7 Flash model card after publication, making changes to language in the “Key Results for Gemini 3.7 Flash” column in the Frontier Safety Assessment section

View details

OpenAI

Aug 3, 2026

OpenAI updated its system card for two of its models: GPT 5.6 and GPT Live.

View details

Anthropic

Jul 24, 2026

Anthropic updated its Frontier Compliance Framework.

View details

xAI

Jul 20, 2026

xAI made changes throughout the model card for Grok 4.5, including some that appear to be persistent errors

View details

Frequently asked questions

Have more questions? Our team is happy to help, contact us.

What does The Midas Project do?
What does your name mean?
Who is behind The Midas Project?
Are you anti-AI?
How can I contribute?
How can I get in touch?