Watchtower

Watchtower tracks when AI companies change their safety policies — and when they break them.

Anthropic

Anthropic’s release of Opus 5.5 with an “inconclusive” determination for harmful-manipulation.

Minor
Concern

Anthropic

Sep 22, 2026

Claude Opus 5.5 was recently released along with its 230-page system card. Within that report, Anthropic’s “harmful manipulation” risk assessment raises the question of what level of certainty is compliant with the company’s Frontier Compliance Framework (FCF). For background, harmful manipulation assesses the model’s “capabilities to conduct influence operations, election interference, or other coordinated campaigns to manipulate public opinion or undermine democratic processes.”1 The FCF provides risk “Tiers” to mark the model’s capabilities. 

Anthropic assessed Opus 5.5 for Tier 2 risks of harmful manipulation but in the system card said the results were “inconclusive.”2 The company says that it does “not believe Claude Opus 5.5 has surpassed the Tier 2 threshold”3 but discloses that Opus 5.5’s scores on the influence-campaign evaluation were “within the range associated with our Tier 2 threshold.”4 Anthropic further notes that this evaluation itself “continues to appear saturated.”5 Overall, Anthropic’s harmful manipulation risk determination rests on the fact that the model’s effectiveness against real people “has not been established.”6 

The FCF does not explicitly require certainty in assignment of risk tiers; however, it does require risk determinations to incorporate appropriate safety margins and generally for Anthropic to document its justification for proceeding.7 Additionally, Anthropic has committed that where its analysis identifies gaps, it will implement and test additional mitigations before deployment.8 

Moreover, when Anthropic found the earlier-released Mythos 5.1 to be similarly “inconclusive” with respect to Tier 2 harmful manipulation, the company explicitly noted that it was developing “a next-generation evaluation designed to more realistically reflect the relevant dynamics.”9 It appears Anthropic released Opus 5.5 before completing that “next-generation evaluation.” Opus 5.5’s system card does not mention the development of that new evaluation.

1. Frontier Compliance Framework, §2.2.

2. Opus 5.5 System Card, p. 79.

3. Opus 5.5 System Card, p. 81.

4. Opus 5.5 System Card, p. 83.

5. Opus 5.5 System Card, p. 79.

6. Opus 5.5 System Card, p. 81.

7. Frontier Compliance Framework, §§2.3, 2.5.

8. Frontier Compliance Framework, §2.2.

9. Claude Fable 5.1 & Claude Mythos 5.1 System Card, p. 81.

OpenAI

Sep 9, 2026

Updated the GPT-6 Astra system card, specifically to hedge claims about misalignment.

View details

OpenAI

Aug 18, 2026

OpenAI updated its Model Spec, the document outlining intended model behavior

View details

xAI

Aug 17, 2026

xAI updated Grok 4.6's model card after release. A changelog was included.

View details

Google

Aug 14, 2026

Google updated its Gemini 3.7 Flash model card after publication, making changes to language in the “Key Results for Gemini 3.7 Flash” column in the Frontier Safety Assessment section

View details

OpenAI

Aug 3, 2026

OpenAI updated its system card for two of its models: GPT 5.6 and GPT Live.

View details

Anthropic

Jul 24, 2026

Anthropic updated its Frontier Compliance Framework.

View details

Frequently asked questions

Have more questions? Our team is happy to help, contact us.

What does The Midas Project do?
What does your name mean?
Who is behind The Midas Project?
Are you anti-AI?
How can I contribute?
How can I get in touch?