Watchtower

Watchtower tracks when AI companies change their safety policies — and when they break them.

xAI

xAI says Grok 4.7 scores below its safety framework’s capability thresholds on dual-use knowledge, but xAI’s safety framework doesn’t provide thresholds.

Moderate
Concern

xAI

Sep 21, 2026

xAI released Grok 4.7, calling it  “our most capable model to date.”

The model card says Grok 4.7 “scores below the FAIF safety thresholds on dual-use knowledge,” and describes the framework as one that “defines… capability thresholds.”1 But the xAI Frontier Artificial Intelligence Framework (FAIF) does not define any thresholds. And notably, the model card only directly covers two of the framework’s four defined risk domains (CBRN and cyber). Loss of control and harmful manipulation are not mentioned by name in the model card.

The FAIF defines harmful manipulation as “models with high manipulative capabilities potentially being misused in ways that could reasonably result in large scale harm.”2 In the Grok 4.7 model card, in the “General output safety” section,3 there’s mention of general refusals but nothing specific to harmful manipulation. 

Loss of control is defined in the FAIF as “risks from humans losing the ability to reliably direct, modify, or shut down a model,”4 and the framework names deception and sycophancy as relevant. The model card reports behavior evaluations: MASK-Rectified, which tests honesty, and sycophancy. These evaluations are described as covering propensities that affect “reliability, neutrality, and behavior.”5 Note that the Grok 4.6 model card used “controllability” instead of “behavior.”6 No evaluation in the card is presented as addressing loss of control.

Per xAI’s framework, releasing a model requires determining that each identified risk is acceptable, a determination the FAIF says incorporates “a margin of security” and draws on estimates of “the probability and severity of harm.”7 The model card never mentions this determination, so it doesn’t say whether or on what basis any of the four risks were judged acceptable.

1. Grok 4.7 Model Card, p. 19.

2. xAI Frontier Artificial Intelligence Framework, p. 1.

3. Grok 4.7 Model Card, p. 23. 

4. xAI Frontier Artificial Intelligence Framework, p. 1.

5. Grok 4.7 Model Card, p. 25.

6. Grok 4.6 Model Card, p. 37.

7. xAI Frontier Artificial Intelligence Framework, pp. 5–6.

Anthropic

Sep 22, 2026

Anthropic’s release of Opus 5.5 with an “inconclusive” determination for harmful-manipulation.

View details

OpenAI

Sep 9, 2026

Updated the GPT-6 Astra system card, specifically to hedge claims about misalignment.

View details

OpenAI

Aug 18, 2026

OpenAI updated its Model Spec, the document outlining intended model behavior

View details

xAI

Aug 17, 2026

xAI updated Grok 4.6's model card after release. A changelog was included.

View details

Google

Aug 14, 2026

Google updated its Gemini 3.7 Flash model card after publication, making changes to language in the “Key Results for Gemini 3.7 Flash” column in the Frontier Safety Assessment section

View details

OpenAI

Aug 3, 2026

OpenAI updated its system card for two of its models: GPT 5.6 and GPT Live.

View details

Frequently asked questions

Have more questions? Our team is happy to help, contact us.

What does The Midas Project do?
What does your name mean?
Who is behind The Midas Project?
Are you anti-AI?
How can I contribute?
How can I get in touch?