Watchtower

Watchtower tracks when AI companies change their safety policies — and when they break them.

OpenAI

Updated the GPT-6 Astra system card, specifically to hedge claims about misalignment.

Moderate
Change

OpenAI

Sep 9, 2026

OpenAI updated its GPT-6 Astra System Card on Sept. 9, 2026—originally published just six days before. The update includes a two paragraph-long change log identifying changes to the Alignment section (Section 8) and changes to the “naming and substance” of the portion dealing with “Verbalized Metagaming and Oversight Gaming” (Section 8.7).

While OpenAI’s Astra release highlighted that the new model represented a step forward in alignment, the revised system card includes new caveats and clarifications suggesting evidence of Astra’s overall alignment should be read in context of the overall system card, including Astra’s higher evaluation awareness and limits of monitoring.

Additionally, the “naming and substance” revisions impact OpenAI’s analysis of its new “honeypot” evaluation—which OpenAI specially crafted following the Hugging Face incident. OpenAI’s system card changes what it previously referred to as “alignment faking” to “oversight gaming” and casts doubt on whether Astra recognized the “honeypot” and thus avoided it, stating that Astra’s metagaming “seems to be largely about whether the task is intentionally impossible.” OpenAI also clarifies that it “did not construct honeypot training environments specifically to target this evaluation or reproduce the incident.”

OpenAI

Aug 18, 2026

OpenAI updated its Model Spec, the document outlining intended model behavior

View details

xAI

Aug 17, 2026

xAI updated Grok 4.6's model card after release. A changelog was included.

View details

Google

Aug 14, 2026

Google updated its Gemini 3.7 Flash model card after publication, making changes to language in the “Key Results for Gemini 3.7 Flash” column in the Frontier Safety Assessment section

View details

OpenAI

Aug 3, 2026

OpenAI updated its system card for two of its models: GPT 5.6 and GPT Live.

View details

Anthropic

Jul 24, 2026

Anthropic updated its Frontier Compliance Framework.

View details

xAI

Jul 20, 2026

xAI made changes throughout the model card for Grok 4.5, including some that appear to be persistent errors

View details

Frequently asked questions

Have more questions? Our team is happy to help, contact us.

What does The Midas Project do?
What does your name mean?
Who is behind The Midas Project?
Are you anti-AI?
How can I contribute?
How can I get in touch?