Watchtower

Watchtower tracks when AI companies change their safety policies — and when they break them.

OpenAI

Seemingly deployed GPT-5.3-Codex without required misalignment safeguards

Moderate
Violation

OpenAI

Feb 5, 2026

OpenAI's GPT-5.3-Codex triggers the "high" cyber risk threshold in OpenAI's own safety framework, which requires specific misalignment safeguards before deployment. OpenAI claims these safeguards aren't needed because the model lacks long-range autonomy (LRA), but the framework's structure contradicts this: one rule requires safeguards for high cyber risk alone, while a separate rule covers any high risk combined with LRA—making the LRA requirement redundant if it applied to both.

OpenAI's previous model already topped METR's benchmark for long-range autonomous task completion, and GPT-5.3-Codex improves on it. The company also acknowledges the model sometimes "sandbags"—deliberately underperforming to avoid triggering safety restrictions—which undermines confidence that LRA is truly absent.

Even if the framework's language were ambiguous, OpenAI should have updated it before deployment rather than relying on a convenient post-hoc interpretation. Under California's SB 53, AI companies must publish a safety framework and adhere to it, with violations carrying fines of $1,000,000 each.

Read more on our Twitter thread: https://x.com/TheMidasProj/status/2019837161647067627

Meta

Oct 2, 2026

Meta updates and renames its safety framework from the "Advanced AI Scaling Framework" (v2) to the "Meta Superintelligence Scaling Framework" (v2.1).

View details

Anthropic

Sep 22, 2026

Anthropic’s release of Opus 5.5 with an “inconclusive” determination for harmful-manipulation.

View details

xAI

Sep 21, 2026

xAI says Grok 4.7 scores below its safety framework’s capability thresholds on dual-use knowledge, but xAI’s safety framework doesn’t provide thresholds.

View details

OpenAI

Sep 9, 2026

Updated the GPT-6 Astra system card, specifically to hedge claims about misalignment.

View details

OpenAI

Aug 18, 2026

OpenAI updated its Model Spec, the document outlining intended model behavior

View details

xAI

Aug 17, 2026

xAI updated Grok 4.6's model card after release. A changelog was included.

View details

Frequently asked questions

Have more questions? Our team is happy to help, contact us.

What does The Midas Project do?
What does your name mean?
Who is behind The Midas Project?
Are you anti-AI?
How can I contribute?
How can I get in touch?