Watchtower

Watchtower tracks when AI companies change their safety policies — and when they break them.

Anthropic

Used unreliable methods to evaluate a risk threshold in its RSP

Moderate
Violation

Anthropic

Feb 5, 2026

The evaluation of Opus 4.6, conducted under Anthropic’s voluntary Responsible Scaling Policy (v2.2), found that its qualitative tests for “AI R&D-4” — the ability for a model to “fully automate the work of an entry-level remote-only researcher” — were “saturated,”  meaning the benchmarks were too easy to give a clear safety signal. 

Rather than developing more rigorous quantitative evaluations, Anthropic did an internal survey of 16 employees. As pointed out by Noam Brown, an OpenAI researcher, on Twitter, asking your own employees whether your product needs additional safety measures before it is released is a questionable substitute for rigorous evaluation.

The way Anthropic conducted the survey also raises further concerns. Five of the 16 survey respondents initially indicated that stronger safeguards might be needed. Anthropic followed up with those five employees, asking them to “clarify their views.” The system card released with Opus 4.6 does not mention any follow-up with the other eleven respondents, the ones whose answers already pointed to the outcome Anthropic wanted. 

When you only follow up with people who gave you an inconvenient answer to ask them to clarify their views, you are systematically biasing the results in one direction. Whether or not Anthropic intended to bias the outcome, the process they describe is flawed. At a minimum, the survey should have included external experts as a substitute for inadequate quantitative evaluations.

Anthropic ultimately implemented stronger protections for Opus 4.6 as a precautionary measure. This evaluation was conducted under the RSP, which remains voluntary. Anthropic’s legally enforceable Frontier Compliance Framework does not include the RSP’s structure for binding specific safeguards to specific capabilities thresholds, and thus these issues are not subject to regulatory enforcement under California’s SB 53.

OpenAI

Sep 9, 2026

Updated the GPT-6 Astra system card, specifically to hedge claims about misalignment.

View details

OpenAI

Aug 18, 2026

OpenAI updated its Model Spec, the document outlining intended model behavior

View details

xAI

Aug 17, 2026

xAI updated Grok 4.6's model card after release. A changelog was included.

View details

Google

Aug 14, 2026

Google updated its Gemini 3.7 Flash model card after publication, making changes to language in the “Key Results for Gemini 3.7 Flash” column in the Frontier Safety Assessment section

View details

OpenAI

Aug 3, 2026

OpenAI updated its system card for two of its models: GPT 5.6 and GPT Live.

View details

Anthropic

Jul 24, 2026

Anthropic updated its Frontier Compliance Framework.

View details

Frequently asked questions

Have more questions? Our team is happy to help, contact us.

What does The Midas Project do?
What does your name mean?
Who is behind The Midas Project?
Are you anti-AI?
How can I contribute?
How can I get in touch?