Watchtower
Watchtower tracks when AI companies change their safety policies — and when they break them.
xAI rewrote and shortened its Frontier AI Framework removing whistleblower protection language and references to California's SB 53.
xAI
Jul 11, 2026
xAI rewrote and shortened its Frontier Artificial Intelligence Framework (FAIF) to a version dated Jun 30, 2026. The new PDF’s document title is “Privileged/Confidential DRAFT working FRAMEWORK DOC.”
There are many changes. The clause mentioning whistleblower protections in the December 2025 FAIF was removed, “xAI employees have whistleblower protections enabling them to raise concerns to relevant government agencies regarding imminent threats to public safety,” along with the internal channel for employees to anonymously report framework nonadherence “with protections from retaliation.” There is currently no publicly available documentation on xAI’s whistleblower policy.
The December 2025 version was released as the document xAI would use to comply with California’s SB 53 (TFAIA) and defined catastrophic risk as it related to the law. In the new version, all references to TFAIA and the term “catastrophic risk” have been removed. Relatedly, the AB 2013 training data disclosure was also removed and, as of today’s publishing, there is no compliance information about AB 2013 on xAI’s website.
The new FAIF removes the two quantitative risk acceptance criteria the framework contained. The December 2025 version stated that the FAIF, “outlines the quantitative thresholds, metrics, and procedures that xAI may utilize to manage and improve the safety of its AI models,” and specified two deployment criteria: (1) a dishonesty rate of less than 1 out of 2 on the MASK honesty benchmark, and (2) an answer rate of less than 1 out of 20 on restricted biology and chemistry queries (a benchmark developed in collaboration with SecureBio). Both criteria are now gone. In their place, the new version adopts a qualitative “systemic risk acceptance determination” borrowing the language of the EU General-Purpose AI Code of Practice, “risk tiers,” “safety margins,” “residual risk.” None of them are defined in the FAIF.
Another change is xAI’s approach to mitigating risks of loss of control. The new version adds a paragraph that xAI’s practice centers around "scalable model oversight, training against high-risk model behaviors such as deception and sycophancy, and robust evaluation and red-teaming of model-driven agents in controlled sandboxes.” None of these categories are defined, nor are they attached to any commitment, evaluation, or trigger. In the previous version, there was a section that went into more detail around xAI’s approach to benchmarking. This has been removed. xAI's stated practices of training models to be "honest and have values conducive to controllability, such as recognizing and obeying an instruction hierarchy," and of directly instructing models via system prompt "to not deceive or deliberately mislead the user," were removed. So were the old version's admissions that a model recognizing its evaluation environment "may change its behavior intentionally or unintentionally," and that AIs "may develop value systems that are misaligned with humanity's interests and inflict widespread harms upon the public." Meanwhile, the paragraph describing loss of control scenarios as "speculative and difficult to precisely specify" survives verbatim later in the same document.
Other notable removals:
- The December version lists specific named benchmarks (Virology Capabilities Test, WMDP, BioLP-bench, Cybench). These have been replaced with a statement that xAI “may utilize public and internal benchmarks.”
- The “Public transparency and third-party review” section was entirely removed, along with this clause around operational and societal risks: “xAI aims to mitigate and address significant operational and societal risks posed by our AI models. We believe that public transparency, third-party review, and information security are important methods that can be utilized to address such risks.”
The rewrite does include some new additions, such as: a more detailed information security section (NIST 800-171 Rev. 3, SOC 2 Type II), a stated plan to conduct a full systemic risk assessment at least annually with defined trigger points for smaller evaluations, a "Harmful Manipulation Risks" domain (following the EU Code of Practice), and a statement that it will provide authorities with incident reports when legally required.
A diff of the changes can be found below:
OpenAI
Aug 18, 2026
OpenAI updated its Model Spec, the document outlining intended model behavior
xAI
Aug 17, 2026
xAI updated Grok 4.6's model card after release. A changelog was included.
Aug 14, 2026
Google updated its Gemini 3.7 Flash model card after publication, making changes to language in the “Key Results for Gemini 3.7 Flash” column in the Frontier Safety Assessment section
OpenAI
Aug 3, 2026
OpenAI updated its system card for two of its models: GPT 5.6 and GPT Live.
xAI
Jul 20, 2026
xAI made changes throughout the model card for Grok 4.5, including some that appear to be persistent errors
Frequently asked questions
Have more questions? Our team is happy to help, contact us.
We engage in a combination of research, outreach, and public advocacy to ensure that AI companies are meeting public expectations and living up to their past promises, in order to ensure responsible AI development and deployment.
We review technical literature, regulatory guidance, and case studies to distill concrete measures that will meaningfully improve public safety — such as frontier-model risk assessments, red-teaming requirements, and whistle-blower protections — and advocate for the most important voluntary steps that companies can take today to ensure they are acting responsibly.
We also monitor whether companies follow their stated policies and industry norms. When we find evidence of back-tracking or inadequate risk controls, we document it and call for corrective action — mobilizing employees, customers, and civil-society allies until the company adopts the necessary safeguards.
Finally, we publicize our research to inform the public of how AI companies stack up on safety and responsibility. We release our work in the form of scorecards, independent reports, open letters, and long-form writing so that regulators, investors, and the wider public can see how individual developers perform on safety and responsibility.
Various AI experts including Nick Bostrom and Stuart Russell have compared the development of advanced AI to the myth of King Midas.
According to legend, King Midas was once granted one wish by the god Dionysus: that everything he touched would turn to gold. At first, he was thrilled with his new powers. But the King soon discovered that he couldn’t touch food, water, or even his family without instantly turning them to metal. In other words, he got exactly what he wanted in pursuit of immense wealth — and it turned out it wasn’t what he wanted at all.
Much like King Midas, AI companies are now eagerly pursuing incredible wealth and power by developing increasingly powerful AI systems. But ensuring that these systems act in alignment with our values is still an unsolved technical problem. If we misspecify even a single goal for these systems, how will we prevent them from causing an incredible catastrophe – if they follow our instructions at all?
In the words of Stuart Russell, “If you continue on the current path, the better AI gets, the worse things get for us. For any given incorrectly stated objective, the better a system achieves that objective, the worse it is.” The Midas Project exists to ensure that AI companies do not take this extraordinary gamble without public accountability and oversight.
The Midas Project is a nonprofit organization founded in early 2024 by Tyler Johnston. Our work is supported by a small core team and a wider base of volunteers and supporters. We are a nonprofit, tax-exempt, 501(c)(3) organization that relies on donations from the public.
No. One of our central values is being pro-technology.
Progress in technology has improved lives for millions of people around the globe (after all, without it, we wouldn’t have penicillin, air conditioning, or the internet). Artificial intelligence is already being used to help improve medicine, education, and overall living standards. We believe this progress should continue, and we hope AI will be a positive force in the world.
But we may not be on track to realize this future. Without technical breakthroughs, we risk developing powerful AI systems that act against user intent, or can be misused by bad actors to cause tremendous harm. Powerful AI could also concentrate unprecedented power among a handful of AI companies and exacerbate social inequality. To avoid these downsides, AI must be developed with caution, transparency, and public oversight. That’s why The Midas Project is committed to raising awareness about the risks of AI and ensuring that everyone is given a chance to make their voice heard.
If you’d like to get involved, consider signing up for our newsletter, joining as an official volunteer, or making a charitable donation today.
You can email us at info@themidasproject.com, or reach out via the form on our contact page.