Watchtower

Watchtower tracks when AI companies change their safety policies — and when they break them.

OpenAI

OpenAI updated its system card for two of its models: GPT 5.6 and GPT Live.

Minor
Change

OpenAI

Aug 3, 2026

Between August 3 and August 4, 2026, OpenAI updated its system card for two of its models: GPT 5.6 and GPT Live.

On August 3, they changed the GPT-5.6 model card to add performance on a new red-teaming approach called GPT-Red. The updated model card has been timestamped and the newly-added content is reflected as new; you can read the full model card here.

On August 4, 2026. OpenAI added the following disclaimer to their model card for GPT Live:

“We identified a configuration mismatch in the safety evaluation setup: the results had been generated using a backend configuration that did not match the final model release. We re-ran the full evaluation using the correct configuration and updated the reported results. Some values remained unchanged after rerunning, while others were corrected. We’ve added an appendix to show the original and corrected values side by side. The overall safety conclusions are unchanged; none of the regressions led to numbers below our safety launch standards.”

The current GPT Live model card, with the added comparison appendix, can be found here.

OpenAI

Sep 9, 2026

Updated the GPT-6 Astra system card, specifically to hedge claims about misalignment.

View details

OpenAI

Aug 18, 2026

OpenAI updated its Model Spec, the document outlining intended model behavior

View details

xAI

Aug 17, 2026

xAI updated Grok 4.6's model card after release. A changelog was included.

View details

Google

Aug 14, 2026

Google updated its Gemini 3.7 Flash model card after publication, making changes to language in the “Key Results for Gemini 3.7 Flash” column in the Frontier Safety Assessment section

View details

Anthropic

Jul 24, 2026

Anthropic updated its Frontier Compliance Framework.

View details

xAI

Jul 20, 2026

xAI made changes throughout the model card for Grok 4.5, including some that appear to be persistent errors

View details

Frequently asked questions

Have more questions? Our team is happy to help, contact us.

What does The Midas Project do?
What does your name mean?
Who is behind The Midas Project?
Are you anti-AI?
How can I contribute?
How can I get in touch?