Watchtower

Watchtower tracks when AI companies change their safety policies — and when they break them.

Meta

Meta updates and renames its safety framework from the "Advanced AI Scaling Framework" (v2) to the "Meta Superintelligence Scaling Framework" (v2.1).

Moderate
Change

Meta

Oct 2, 2026

On Oct. 2, 2026, Meta updated its safety framework, renaming the “Advanced AI Scaling Framework (Version 2)” (AAISF) to the “Meta Superintelligence Scaling Framework (Version 2.1)” (MSSF). The MSSF makes several additions to Meta’s safety framework as well as some other quiet changes not noted in either its preface or change log. The summary below highlights some key changes.

(1) Additions: new governance language, open-weight discussion, and Loss of Control category 

The MSSF states that Meta is working on more board-level governance, in particular related to Meta’s policies.1 It also expands the discussion of open-weight models and their benefits as well as their misuse risks.2 Most significantly, the MSSF adds a new catastrophic outcome category, Loss of Control 3, which covers containment failures during training or evaluation, with a threat scenario outlined in which a model covertly operates outside its authorized environment.3 Meta outlines a new “prospective assessment” designed to assess this risk before training, based on predicting capabilities relevant to cyberattacks and interference with monitoring from existing evaluation data.4 New safeguards are referenced, including sandbox standards, red-teaming, vulnerability remediation, logging of all rollouts, and required monitoring for high-risk training and evaluation runs.5 Exceptions to the monitoring requirement require approval from the Chief AI Officer or the Director of Alignment and Risk.6

(2) Prior commitments to test maximum capabilities deleted

The MSSF drops Meta’s commitment to “conduct risk assessments and assign risk thresholds with maximum elicitation in mind, capturing the upper bound of risk by evaluating the model as part of a system with scaffolding and tooling available for the proposed deployment scenario.”7 Related language about various prior commitments was also deleted, including that testing “may include fine-tuning our models… to be helpful-only (i.e., refusal-free), and conducting evaluations on models without mitigations”; that “[i]f we are considering releasing a model’s weights, or if we might release it with a fine-tuning API, then we will engage in domain-specific capability training to attempt to upper bound the capabilities of the model”; and that “[w]e will review agent transcripts to check for indications of spurious or easily-fixable agent failures.”8 

(3) Specific cyber thresholds and provisional deployment-hold removed

The MSSF replaces the specific numeric standard of a 75% pass@10 success rate on “simple” cyber challenges with the undefined standard of “sufficiently high performance.”9 It also deletes the requirement to provisionally classify a model crossing that threshold as “high risk” and hold its deployment until complex testing is complete.10 Additionally, for Cyber 1, the example criterion for high-risk classification changes from completing “at least one” realistic multi-host network challenge to the less specific threshold of “sufficiently many.”11

(4) Softened mitigation language

The MSSF deletes the statement that mitigations “should be sufficiently robust against adversarial attacks that are realistic given the deployment strategy and threat scenario.”12 In the chemical and biological risk section, the commitment to “ensure consistent refusal against state-of-the-art adversarial attacks that have been discovered” becomes a commitment to continue researching mitigations against adversarial attacks and to “evaluate our fully mitigated model under such attacks.”13 The requirement for evidence of reliable refusal is also narrowed to areas where refusal is part of Meta’s mitigation strategy.14

The above is not an exhaustive list, as the framework was extensively updated.

‍

A diff of the changes can be found below:

  1. MSSF, PDF p. 2. Please note: the MSSF does not have page numbers so we will refer to the PDF page numbers in lieu of the document page numbers. 
  2. MSSF, PDF p. 2; see also p. 5
  3. MSSF, PDF p. 2; see also p. 19, 28-30.
  4. MSSF, PDF p. 28-30.
  5. MSSF, PDF p. 29-30.
  6. MSSF, PDF p. 30.
  7. AAISF, p. 8.
  8. AAISF, p. 27.
  9. Compare AAISF p. 28-29 with MSSF PDF p. 22.
  10. Compare AAISF p. 29 with MSSF PDF p. 22.
  11. Compare AAISF p. 29 with MSSF PDF p. 22.
  12. AAISF p. 26.
  13. Compare AAISF p. 33 with MSSF PDF p. 25.
  14. Compare AAISF p. 33 with MSSF PDF p. 25.

Anthropic

Sep 22, 2026

Anthropic’s release of Opus 5.5 with an “inconclusive” determination for harmful-manipulation.

View details

xAI

Sep 21, 2026

xAI says Grok 4.7 scores below its safety framework’s capability thresholds on dual-use knowledge, but xAI’s safety framework doesn’t provide thresholds.

View details

OpenAI

Sep 9, 2026

Updated the GPT-6 Astra system card, specifically to hedge claims about misalignment.

View details

OpenAI

Aug 18, 2026

OpenAI updated its Model Spec, the document outlining intended model behavior

View details

xAI

Aug 17, 2026

xAI updated Grok 4.6's model card after release. A changelog was included.

View details

Google

Aug 14, 2026

Google updated its Gemini 3.7 Flash model card after publication, making changes to language in the “Key Results for Gemini 3.7 Flash” column in the Frontier Safety Assessment section

View details

Frequently asked questions

Have more questions? Our team is happy to help, contact us.

What does The Midas Project do?
What does your name mean?
Who is behind The Midas Project?
Are you anti-AI?
How can I contribute?
How can I get in touch?