Watchtower
Watchtower tracks when AI companies change their safety policies — and when they break them. Click any entry for further details.
OpenAI
Aug 18, 2026
OpenAI updated its Model Spec, the document outlining intended model behavior
xAI
Aug 17, 2026
xAI updated Grok 4.6's model card after release. A changelog was included.
Aug 14, 2026
Google updated its Gemini 3.7 Flash model card after publication, making changes to language in the “Key Results for Gemini 3.7 Flash” column in the Frontier Safety Assessment section
OpenAI
Aug 3, 2026
OpenAI updated its system card for two of its models: GPT 5.6 and GPT Live.
Anthropic
Jul 24, 2026
Anthropic updated its Frontier Compliance Framework.
xAI
Jul 20, 2026
xAI made changes throughout the model card for Grok 4.5, including some that appear to be persistent errors
xAI
Jul 11, 2026
xAI rewrote and shortened its Frontier AI Framework removing whistleblower protection language and references to California's SB 53.
Anthropic
Jul 8, 2026
Anthropic updated its Responsible Scaling Policy (RSP) from v3.3 to v3.4. Including changes to the Automated R&D "dramatic acceleration" trigger and adjusting Risk
Anthropic
Jun 9, 2026
Anthropic released a white paper describing its practices around retention and review of enterprise customer data.
Anthropic
Jun 8, 2026
Anthropic updated its Frontier Compliance Framework to match language in its RSP and expanded the “Sabotage and loss of control” Tier 2.
OpenAI
May 29, 2026
OpenAI published a separate Frontier Governance Framework (FGF) to satisfy California’s Transparency in Frontier AI Act and the EU Code of Practice.
Anthropic
May 26, 2026
Anthropic updated its Responsible Scaling Policy (RSP) from v3.2 to v3.3.
Anthropic
Apr 29, 2026
Anthropic made minor updates to the RSP expanding on the Long Term Benefit Trust's formal role.
Apr 17, 2026
Google updated its Frontier Safety Framework from version 3.0 to 3.1, in a change announced on its website.
Meta
Apr 7, 2026
Meta released v2 of its AI safety policy, now called the "Advanced AI Scaling Framework," substantially rewriting and expanding the original with more granular
Anthropic
Apr 7, 2026
Slight updates to the Claude Mythos Preview System Card
Anthropic
Apr 2, 2026
Anthropic made minor updates to its RSP, clarifying its AI R&D automation threshold and its ability to take precautionary actions beyond what the policy requires
Anthropic
Mar 24, 2026
Updated its RSP noncompliance reporting and anti-retaliation policy
Anthropic
Mar 2, 2026
Updated its Frontier Compliance Framework without public announcement
Anthropic
Feb 24, 2026
Revoked a core component of their Responsible Scaling Policy
Feb 12, 2026
Google released a new model without a required safety scorecard or clear communication about the nature of the release
OpenAI
Feb 5, 2026
Seemingly deployed GPT-5.3-Codex without required misalignment safeguards
Anthropic
Feb 5, 2026
Used unreliable methods to evaluate a risk threshold in its RSP
Anthropic
Dec 19, 2025
Published a secondary compliance framework, substituting the RSP for the sake of new compliance requirements from the EU and California
xAI
Aug 28, 2025
xAI released Grok Code Fast 1 despite the model failing a safety test that the company's policy says models must pass before release.
xAI
Aug 22, 2025
Released Grok 4 System Card and new changes to RMF, including immediate quiet redactions
xAI
Jul 9, 2025
xAI launched Grok 4 without a safety report, violating its Seoul commitment.
Anthropic
May 14, 2025
Anthropic removed its commitment to define ASL-4 evaluations before reaching ASL-3, and released an ASL-3 model without some promised safeguards.
OpenAI
Apr 14, 2025
OpenAI did not release a safety scorecard for GPT-4.1 despite promises to do so.
Anthropic
Mar 31, 2025
Released v2.1 of their Responsible Scaling Policy
Mar 25, 2025
Google initially released Gemini 2.5 Pro without a safety report, in violation of a commitment to the White House.
Mar 6, 2025
Scrubbed mentions of diversity and equity from the mission description of their Responsible AI team.
Anthropic
Feb 27, 2025
Removed "White House's Voluntary Commitments for Safe, Secure, and Trustworthy AI"
Feb 4, 2025
On February 4th, Google released a new version of their Frontier Safety Framework.
Feb 4, 2025
Removed previous commitment not to develop AI for use in warfare or surveillance
Meta
Feb 3, 2025
Released a responsible scaling policy, entitled their Frontier AI Framework.
OpenAI
Jan 17, 2025
Made substantial changes throughout the o1 system card. Did not announce these changes.
OpenAI
Jan 14, 2025
Adjusted the language on the o1 system card webpage, changing "o1" to "o1-preview."
Microsoft
Dec 23, 2024
Removed Vice Chair and President Brad Smith's byline from Microsoft's 2023 White House commitment to advance safe and secure artificial intelligence.
Anthropic
Dec 19, 2024
Between December 16 and December 18, Anthropic changed the "last updated" date on their Responsible Disclosure Policy, with no apparent substantive changes to the text of the policy.

Cognition
Dec 11, 2024
Changed terms of service concerning use of user data. Did not announce or report that change was made.
Cohere
Nov 21, 2024
On November 21, Cohere released a complete rewrite of their usage policies.
OpenAI
Nov 21, 2024
Released a white paper detailing how they approach external red teaming.
Anthropic
Nov 18, 2024
Released a new page providing details about how they are complying with multiple voluntary safety and security frameworks.
Anthropic
Oct 15, 2024
Released an updated version of their Responsible Scaling Policy.
Magic.dev
Sep 18, 2024
Released a statement on AI safety priorities, and announced an upcoming v2 of their responsible scaling policy.
OpenAI
Sep 12, 2024
Released preparedness scorecard for their newest model, o1.

Cognition
Sep 4, 2024
Released an acceptable usage policy, along with a reporting email for security vulnerabilities.
OpenAI
Aug 18, 2024
Adjusted authorship for a two-year-old article on their approach to alignment (with no substantive changes to the content)
OpenAI
Aug 8, 2024
Released the preparedness scorecard for GPT-4o (many months behind promised schedule)
OpenAI
Jul 4, 2024
OpenAI had a major cybersecurity incident and failed to report it for over a year.
Magic.dev
Jul 2, 2024
Released a responsible scaling policy, entitled their "AGI Readiness Policy."
Meta
Jun 26, 2024
Updated privacy policy to permit the use of personal user information (photos, posts, etc.) to train Meta AI models.
Cohere
May 24, 2024
Changed commitment for access reviews from "quarterly" to "periodic"
OpenAI
May 21, 2024
Reported to have abandoned former promise to dedicate 20% of compute resources to advanced AI alignment.
OpenAI
May 21, 2024
OpenAI never fulfilled its promise to allocate resources to its own safety team.
May 17, 2024
Released Frontier Safety Framework, Google's response to Anthropic's RSP and OpenAI's Preparedness Framework.
OpenAI
Apr 16, 2024
OpenAI waited to release a promised safety evaluation until three months after the model had already been publicly released.
OpenAI
Jan 24, 2024
Quietly scrapped policies allowing public inspection of governance documents, financial statements, and conflict of interest rules.
OpenAI
Jan 10, 2024
Changed their usage policies to remove a ban on using OpenAI products for "military and warfare."
OpenAI
Oct 12, 2023
Quietly changed "core values" to de-emphasize social impact and emphasize an intense focus on AGI.
OpenAI
Dec 1, 2022
OpenAI’s GPT-4 was publicly tested in India without the required approval from its Development Safety Board.
There are no watchtowers based on these filters. Please try searching again.
Frequently asked questions
Have more questions? Our team is happy to help, contact us.
We engage in a combination of research, outreach, and public advocacy to ensure that AI companies are meeting public expectations and living up to their past promises, in order to ensure responsible AI development and deployment.
We review technical literature, regulatory guidance, and case studies to distill concrete measures that will meaningfully improve public safety — such as frontier-model risk assessments, red-teaming requirements, and whistle-blower protections — and advocate for the most important voluntary steps that companies can take today to ensure they are acting responsibly.
We also monitor whether companies follow their stated policies and industry norms. When we find evidence of back-tracking or inadequate risk controls, we document it and call for corrective action — mobilizing employees, customers, and civil-society allies until the company adopts the necessary safeguards.
Finally, we publicize our research to inform the public of how AI companies stack up on safety and responsibility. We release our work in the form of scorecards, independent reports, open letters, and long-form writing so that regulators, investors, and the wider public can see how individual developers perform on safety and responsibility.
Various AI experts including Nick Bostrom and Stuart Russell have compared the development of advanced AI to the myth of King Midas.
According to legend, King Midas was once granted one wish by the god Dionysus: that everything he touched would turn to gold. At first, he was thrilled with his new powers. But the King soon discovered that he couldn’t touch food, water, or even his family without instantly turning them to metal. In other words, he got exactly what he wanted in pursuit of immense wealth — and it turned out it wasn’t what he wanted at all.
Much like King Midas, AI companies are now eagerly pursuing incredible wealth and power by developing increasingly powerful AI systems. But ensuring that these systems act in alignment with our values is still an unsolved technical problem. If we misspecify even a single goal for these systems, how will we prevent them from causing an incredible catastrophe – if they follow our instructions at all?
In the words of Stuart Russell, “If you continue on the current path, the better AI gets, the worse things get for us. For any given incorrectly stated objective, the better a system achieves that objective, the worse it is.” The Midas Project exists to ensure that AI companies do not take this extraordinary gamble without public accountability and oversight.
The Midas Project is a nonprofit organization founded in early 2024 by Tyler Johnston. Our work is supported by a small core team and a wider base of volunteers and supporters. We are a nonprofit, tax-exempt, 501(c)(3) organization that relies on donations from the public.
No. One of our central values is being pro-technology.
Progress in technology has improved lives for millions of people around the globe (after all, without it, we wouldn’t have penicillin, air conditioning, or the internet). Artificial intelligence is already being used to help improve medicine, education, and overall living standards. We believe this progress should continue, and we hope AI will be a positive force in the world.
But we may not be on track to realize this future. Without technical breakthroughs, we risk developing powerful AI systems that act against user intent, or can be misused by bad actors to cause tremendous harm. Powerful AI could also concentrate unprecedented power among a handful of AI companies and exacerbate social inequality. To avoid these downsides, AI must be developed with caution, transparency, and public oversight. That’s why The Midas Project is committed to raising awareness about the risks of AI and ensuring that everyone is given a chance to make their voice heard.
If you’d like to get involved, consider signing up for our newsletter, joining as an official volunteer, or making a charitable donation today.
You can email us at info@themidasproject.com, or reach out via the form on our contact page.
