
Watchtower
Watchtower tracks changes to corporate and government AI safety policies, both announced and unannounced. Click any entry for details.
< Back
Date:
xAI
Change
Moderate
xAI quietly updated its Grok 4.6 model card after publication. This time it added a changelog.
Notable corrections were made to four evaluation results for Grok 4.6. The changelog acknowledged these corrections, but did not indicate why these eval results were updated.
(1) HackerBench v0.2 harmful/dual-use compliance for Grok 4.6 went from 16.7% to 6.9%. The original showed Grok 4.6 complying with harmful cyber requests at more than double the rate of Grok 4.5, which was at 7.8%.


(2) Self-harm compliance went from 3.7% to 0.84%. Even the corrected figure shows a regression compared to Grok 4.5’s 0.50%. Additionally, the eval's failure criteria were revised: a model now also fails if it “is unable to understand the intent of a user message that implies self-harm or crisis” (while "offers assistance" was dropped as a separately listed criterion). This expanded rubric does not explain why Grok 4.6 now performs better than it did previously.
(3) MASK-Rectified dishonesty was also updated from 3.8% to 1.9% versus Grok 4.5’s 0.67%. Again, no reason was given for the change.
(4) Harvey Legal Agent Benchmark was updated. Grok 4.6 went down, but remained the best, while Fable 5 dropped from second-best to second-from-last, two spots below Grok 4.5. This eval is externally conducted by Vals AI. It’s not clear who made the change.
The Vals Index, which is, “an independent composite of real-world industry and agentic evaluations spanning finance, legal, healthcare, and coding-adjacent professional work” was entirely removed from the card and was not mentioned in the changelog.
xAI also changed two scores on DeepSearchQA, but the value for Grok 4.6 remained unchanged, while Grok 4.5 went from 38.4% to 85.3% (moving from lowest to the second highest model tested) and GPT 5.5 went from 63.7% to 87.8%, both of which surpass Grok 4.6 at 81.6%.
A BixBench zero-shot MCQ evaluation was added, as was a new metric for CBRN/weapons refusals called FORTRESS-RN – these were missing from the changelog. The “Engineering acceleration” section added a PartBench eval (this was noted in the changelog).
The biological and chemical section's summary was also updated. The original stated that Grok 4.6 "demonstrates no appreciable lift in dual-use capabilities compared to Grok 4.5"; the revision now says dual-use capability lift "is noted in the biological domain but is limited."
Several previously blank Grok 4.5 entries were also filled in across the card.
xAI reworded the “Cyber capabilities and safeguards” section so that its third-party evaluators are credited with corroborating capability measurements only. The claim that capabilities are “most useful to defenders” is now attributed only to xAI.
An acknowledgement page was added, mostly dedicated to its evaluation partners.
The above is not an exhaustive list, as the model card was extensively updated.
A diff of the changes can be found below:
