
Watchtower
Watchtower tracks changes to corporate and government AI safety policies, both announced and unannounced. Click any entry for details.
< Back
Date:
xAI
Change
Moderate
Around July 20, 2026, xAI quietly updated its model card for Grok 4.5, adding a small “Revision: 2026-07-20” timestamp at the bottom of the document’s title page.
Various minor changes were made throughout the model card, and an entire new section on “Engineering acceleration” was added. Among the changes were small revisions to the reporting of model performance on benchmarks. For example, results that were previously attributed to Opus 4.8 and Sonnet 5 on “max” reasoning (and which showed Grok 4.5 outperforming both models) were revised to instead say “high” reasoning. Somewhat concerningly, the bio/chem section seemed to have initially misreported how performance was measured: the original section asserted for one benchmark that the model was scored after applying safeguards, whereas the new version reads: “Unless otherwise stated, the benchmarks presented in Section 8 are performed without any safeguards in place.” Both versions still present the same score. Additionally, one of the revised benchmarks may contain a copy-pasting error; on DeepSearchQA, Opus 4.8 and Grok 4.5 are scored at 40.7% and 38.4%. This is surprisingly low compared to current SOTA performance on the benchmark, as reported by other parties. As it turns out, the two numbers are identical to the same models’ reported performance on RelBench in the previous section, suggesting that both models’ performance on the benchmark is misreported.
A diff of the changes can be found below:
