Black lighthouse silhouette graphic – visual asset for The Midas Project Watchtower page.

Watchtower

Watchtower tracks changes to corporate and government AI safety policies, both announced and unannounced. Click any entry for details.

< Back

Date:

Anthropic

Change

Moderate

Anthropic updated its Frontier Compliance Framework (FCF), the compliance-facing document that serves as Anthropic's framework under California's TFAIA and the EU General-Purpose AI Code of Practice (EU CoP). Most notably, the criteria for activating the automated AI R&D risk tier were updated to match Anthropic’s current Responsible Scaling Policy (v 3.4). For the tier to be activated, the FCF now requires a doubling in the rate of AI progress relative to both the expected rate of progress and the fastest rate of extended progress previously observed in the absence of significant AI contributions. This continued trend must also “seem likely” to lead to greater acceleration in capabilities. (See further analysis of these changes in our update on Anthropic RSP v. 3.4.) 

The new FCF also includes harmful manipulation as one of the areas subject to pre-launch risk analysis, alongside CBRN, loss of control, and cyber offense risks. For the “Misaligned AI systems in high-stakes settings” risk tier (Tier 1 under Loss of Control), AI systems writing “large amounts of critical code” is no longer specifically named as a factor that could activate this risk tier. 

A diff of the changes can be found below: