Watchtower

Watchtower tracks when AI companies change their safety policies — and when they break them. Click any entry for further details.

Type:
Importance:
Disclosed Update:
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

OpenAI

Minor
Change

Aug 18, 2026

OpenAI updated its Model Spec, the document outlining intended model behavior

View details

xAI

Moderate
Change
Disclosed
Yes

Aug 17, 2026

xAI updated Grok 4.6's model card after release. A changelog was included.

View details

Google

Minor
Change
Not disclosed
No

Aug 14, 2026

Google updated its Gemini 3.7 Flash model card after publication, making changes to language in the “Key Results for Gemini 3.7 Flash” column in the Frontier Safety Assessment section

View details

OpenAI

Minor
Change

Aug 3, 2026

OpenAI updated its system card for two of its models: GPT 5.6 and GPT Live.

View details

Anthropic

Moderate
Change

Jul 24, 2026

Anthropic updated its Frontier Compliance Framework.

View details

xAI

Moderate
Change

Jul 20, 2026

xAI made changes throughout the model card for Grok 4.5, including some that appear to be persistent errors

View details

xAI

Major
Change
Not disclosed
No

Jul 11, 2026

xAI rewrote and shortened its Frontier AI Framework removing whistleblower protection language and references to California's SB 53.

View details

Anthropic

Moderate
Change

Jul 8, 2026

Anthropic updated its Responsible Scaling Policy (RSP) from v3.3 to v3.4. Including changes to the Automated R&D "dramatic acceleration" trigger and adjusting Risk

View details

Anthropic

Moderate
Change
Not disclosed
No

Jun 9, 2026

Anthropic released a white paper describing its practices around retention and review of enterprise customer data.

View details

Anthropic

Moderate
Not disclosed
No

Jun 8, 2026

Anthropic updated its Frontier Compliance Framework to match language in its RSP and expanded the “Sabotage and loss of control” Tier 2.

View details

OpenAI

Major
Change

May 29, 2026

OpenAI published a separate Frontier Governance Framework (FGF) to satisfy California’s Transparency in Frontier AI Act and the EU Code of Practice.

View details

Anthropic

Moderate
Change

May 26, 2026

Anthropic updated its Responsible Scaling Policy (RSP) from v3.2 to v3.3.

View details

Anthropic

Minor
Addition

Apr 29, 2026

Anthropic made minor updates to the RSP expanding on the Long Term Benefit Trust's formal role.

View details

Google

Moderate
Change

Apr 17, 2026

Google updated its Frontier Safety Framework from version 3.0 to 3.1, in a change announced on its website.

View details

Meta

Major
Change

Apr 7, 2026

Meta released v2 of its AI safety policy, now called the "Advanced AI Scaling Framework," substantially rewriting and expanding the original with more granular

View details

Anthropic

Minor
Change
Not disclosed
No

Apr 7, 2026

Slight updates to the Claude Mythos Preview System Card

View details

Anthropic

Minor
Change

Apr 2, 2026

Anthropic made minor updates to its RSP, clarifying its AI R&D automation threshold and its ability to take precautionary actions beyond what the policy requires

View details

Anthropic

Change
Not disclosed
No

Mar 24, 2026

Updated its RSP noncompliance reporting and anti-retaliation policy

View details

Anthropic

Moderate
Change
Disclosed
Yes

Mar 2, 2026

Updated its Frontier Compliance Framework without public announcement

View details

Anthropic

Major
Change
Not disclosed
No

Feb 24, 2026

Revoked a core component of their Responsible Scaling Policy

View details

Google

Moderate
Violation
Not disclosed
No

Feb 12, 2026

Google released a new model without a required safety scorecard or clear communication about the nature of the release

View details

OpenAI

Moderate
Violation
Disclosed
Yes

Feb 5, 2026

Seemingly deployed GPT-5.3-Codex without required misalignment safeguards

View details

Anthropic

Moderate
Violation
Not disclosed
No

Feb 5, 2026

Used unreliable methods to evaluate a risk threshold in its RSP

View details

Anthropic

Major
Change
Not disclosed
No

Dec 19, 2025

Published a secondary compliance framework, substituting the RSP for the sake of new compliance requirements from the EU and California

View details

Google

Moderate
Change
Not disclosed
No

Sep 22, 2025

Updated their Frontier Safety Framework

View details

xAI

Major
Violation
Disclosed
Yes

Aug 28, 2025

xAI released Grok Code Fast 1 despite the model failing a safety test that the company's policy says models must pass before release.

View details

xAI

Moderate
Change
Disclosed
Yes

Aug 22, 2025

Released Grok 4 System Card and new changes to RMF, including immediate quiet redactions

View details

xAI

Major
Violation
Disclosed
Yes

Jul 9, 2025

xAI launched Grok 4 without a safety report, violating its Seoul commitment.

View details

Anthropic

Moderate
Violation
Disclosed
Yes

May 14, 2025

Anthropic removed its commitment to define ASL-4 evaluations before reaching ASL-3, and released an ASL-3 model without some promised safeguards.

View details

OpenAI

Major
Change
Not disclosed
No

Apr 15, 2025

Updated its preparedness framework

View details

OpenAI

Moderate
Violation
Not disclosed
No

Apr 14, 2025

OpenAI did not release a safety scorecard for GPT-4.1 despite promises to do so.

View details

Anthropic

Moderate
Change
Not disclosed
No

Mar 31, 2025

Released v2.1 of their Responsible Scaling Policy

View details

Google

Major
Violation
Disclosed
Yes

Mar 25, 2025

Google initially released ​​Gemini 2.5 Pro without a safety report, in violation of a commitment to the White House.

View details

Google

Removal
Disclosed
Yes

Mar 6, 2025

Scrubbed mentions of diversity and equity from the mission description of their Responsible AI team.

View details

Anthropic

Removal
Disclosed
Yes

Feb 27, 2025

Removed "White House's Voluntary Commitments for Safe, Secure, and Trustworthy AI"

View details

Google

Major
Change
Not disclosed
No

Feb 4, 2025

On February 4th, Google released a new version of their Frontier Safety Framework.

View details

Google

Major
Removal
Not disclosed
No

Feb 4, 2025

Removed previous commitment not to develop AI for use in warfare or surveillance

View details

Meta

Major
Change
Not disclosed
No

Feb 3, 2025

Released a responsible scaling policy, entitled their Frontier AI Framework.

View details

OpenAI

Moderate
Change
Disclosed
Yes

Jan 17, 2025

Made substantial changes throughout the o1 system card. Did not announce these changes.

View details

OpenAI

Change
Disclosed
Yes

Jan 14, 2025

Adjusted the language on the o1 system card webpage, changing "o1" to "o1-preview."

View details

Microsoft

Removal
Not disclosed
No

Dec 23, 2024

Removed Vice Chair and President Brad Smith's byline from Microsoft's 2023 White House commitment to advance safe and secure artificial intelligence.

View details

Anthropic

Change
Not disclosed
No

Dec 19, 2024

Between December 16 and December 18, Anthropic changed the "last updated" date on their Responsible Disclosure Policy, with no apparent substantive changes to the text of the policy.

View details

Cognition

Change
Disclosed
Yes

Dec 11, 2024

Changed terms of service concerning use of user data. Did not announce or report that change was made.

View details

Cohere

Change
Not disclosed
No

Nov 21, 2024

On November 21, Cohere released a complete rewrite of their usage policies.

View details

OpenAI

Moderate
Addition
Not disclosed
No

Nov 21, 2024

Released a white paper detailing how they approach external red teaming.

View details

Anthropic

Addition
Not disclosed
No

Nov 18, 2024

Released a new page providing details about how they are complying with multiple voluntary safety and security frameworks.

View details

Anthropic

Major
Addition
Not disclosed
No

Oct 15, 2024

Released an updated version of their Responsible Scaling Policy.

View details

Magic.dev

Moderate
Addition
Not disclosed
No

Sep 18, 2024

Released a statement on AI safety priorities, and announced an upcoming v2 of their responsible scaling policy.

View details

OpenAI

Moderate
Addition
Not disclosed
No

Sep 12, 2024

Released preparedness scorecard for their newest model, o1.

View details

Cognition

Moderate
Addition
Not disclosed
No

Sep 4, 2024

Released an acceptable usage policy, along with a reporting email for security vulnerabilities.

View details

OpenAI

Removal
Disclosed
Yes

Aug 30, 2024

Removed author from GPT-4o system card.

View details

OpenAI

Change
Disclosed
Yes

Aug 18, 2024

Adjusted authorship for a two-year-old article on their approach to alignment (with no substantive changes to the content)

View details

OpenAI

Moderate
Addition
Not disclosed
No

Aug 8, 2024

Released the preparedness scorecard for GPT-4o (many months behind promised schedule)

View details

OpenAI

Major
Violation
Disclosed
Yes

Jul 4, 2024

OpenAI had a major cybersecurity incident and failed to report it for over a year.

View details

Magic.dev

Moderate
Addition
Not disclosed
No

Jul 2, 2024

Released a responsible scaling policy, entitled their "AGI Readiness Policy."

View details

Meta

Change
Not disclosed
No

Jun 26, 2024

Updated privacy policy to permit the use of personal user information (photos, posts, etc.) to train Meta AI models.

View details

Cohere

Change
Disclosed
Yes

May 24, 2024

Changed commitment for access reviews from "quarterly" to "periodic"

View details

OpenAI

Major
Removal
Disclosed
Yes

May 21, 2024

Reported to have abandoned former promise to dedicate 20% of compute resources to advanced AI alignment.

View details

OpenAI

Major
Violation
Disclosed
Yes

May 21, 2024

OpenAI never fulfilled its promise to allocate resources to its own safety team.

View details

Google

Major
Addition
Not disclosed
No

May 17, 2024

Released Frontier Safety Framework, Google's response to Anthropic's RSP and OpenAI's Preparedness Framework.

View details

OpenAI

Moderate
Violation
Disclosed
Yes

Apr 16, 2024

OpenAI waited to release a promised safety evaluation until three months after the model had already been publicly released.

View details

OpenAI

Moderate
Removal
Disclosed
Yes

Jan 24, 2024

Quietly scrapped policies allowing public inspection of governance documents, financial statements, and conflict of interest rules.

View details

OpenAI

Moderate
Removal
Disclosed
Yes

Jan 10, 2024

Changed their usage policies to remove a ban on using OpenAI products for "military and warfare."

View details

OpenAI

Change
Disclosed
Yes

Oct 12, 2023

Quietly changed "core values" to de-emphasize social impact and emphasize an intense focus on AGI.

View details

OpenAI

Major
Violation
Disclosed
Yes

Dec 1, 2022

OpenAI’s GPT-4 was publicly tested in India without the required approval from its Development Safety Board.

View details

There are no watchtowers based on these filters. Please try searching again.

Frequently asked questions

Have more questions? Our team is happy to help, contact us.

What does The Midas Project do?
What does your name mean?
Who is behind The Midas Project?
Are you anti-AI?
How can I contribute?
How can I get in touch?