Watchtower

Watchtower tracks when AI companies change their safety policies — and when they break them.

xAI

xAI made changes throughout the model card for Grok 4.5, including some that appear to be persistent errors

Moderate
Change

xAI

Jul 20, 2026

Around July 20, 2026, xAI quietly updated its model card for Grok 4.5, adding a small “Revision: 2026-07-20” timestamp at the bottom of the document’s title page.

Various minor changes were made throughout the model card, and an entire new section on “Engineering acceleration” was added. Among the changes were small revisions to the reporting of model performance on benchmarks. For example, results that were previously attributed to Opus 4.8 and Sonnet 5 on “max” reasoning (and which showed Grok 4.5 outperforming both models) were revised to instead say “high” reasoning. Somewhat concerningly, the bio/chem section seemed to have initially misreported how performance was measured: the original section asserted for one benchmark that the model was scored after applying safeguards, whereas the new version reads: “Unless otherwise stated, the benchmarks presented in Section 8 are performed without any safeguards in place.” Both versions still present the same score. Additionally, one of the revised benchmarks may contain a copy-pasting error; on DeepSearchQA, Opus 4.8 and Grok 4.5 are scored at 40.7% and 38.4%. This is surprisingly low compared to current SOTA performance on the benchmark, as reported by other parties. As it turns out, the two numbers are identical to the same models’ reported performance on RelBench in the previous section, suggesting that both models’ performance on the benchmark is misreported.

A diff of the changes can be found below:

OpenAI

Sep 9, 2026

Updated the GPT-6 Astra system card, specifically to hedge claims about misalignment.

View details

OpenAI

Aug 18, 2026

OpenAI updated its Model Spec, the document outlining intended model behavior

View details

xAI

Aug 17, 2026

xAI updated Grok 4.6's model card after release. A changelog was included.

View details

Google

Aug 14, 2026

Google updated its Gemini 3.7 Flash model card after publication, making changes to language in the “Key Results for Gemini 3.7 Flash” column in the Frontier Safety Assessment section

View details

OpenAI

Aug 3, 2026

OpenAI updated its system card for two of its models: GPT 5.6 and GPT Live.

View details

Anthropic

Jul 24, 2026

Anthropic updated its Frontier Compliance Framework.

View details

Frequently asked questions

Have more questions? Our team is happy to help, contact us.

What does The Midas Project do?
What does your name mean?
Who is behind The Midas Project?
Are you anti-AI?
How can I contribute?
How can I get in touch?