The White House's Secret AI Evaluation Framework: What Product Teams Must Do Now
TL;DR
On August 3-4, 2026, the White House finalized its AI model evaluation framework under EO 14409 and immediately refused to release it publicly. The benchmarks and thresholds are classified. The NSA plays a central role in assessments. Open-weight models are excluded entirely. And the 30-day pre-release window for frontier models with advanced cybersecurity capabilities is now operational. For most AI product teams, this framework is not yet your concern: it targets frontier model labs, not feature builders. But if your product is approaching frontier capability in cybersecurity or CBRN domains, the compliance window is already running.
The AI PM Minute
One tactic to make you a sharper AI PM, twice a week. 60 seconds to read. Free.
No fluff. Unsubscribe anytime.
What Happened and When
Executive Order 14409, signed June 2, 2026, gave the government until August 1 to finalize a framework for evaluating frontier AI models before public release. The White House met that deadline: on August 3, the framework was finalized. On August 4, the administration held a staff-level briefing with OpenAI, Google, Anthropic, Meta, Nvidia, and Microsoft. And then it refused to release the framework publicly.
The framework itself is unclassified. The benchmarks and model capability thresholds it establishes are classified. The administration's stated rationale: publicly disclosing the specific thresholds would give adversaries a roadmap for calibrating models to just below the review trigger. Companies and researchers not included in the briefing are left to infer the framework's content from what they know about EO 14409's stated intent.
June 2, 2026
EO 14409 signed. Sets August 1 deadline for frontier AI evaluation framework.
Late July 2026
OpenAI, Google, and Anthropic reviewed a draft framework and submitted edits.
August 1, 2026
White House meets its own deadline, finalizing the framework.
August 3, 2026
Framework finalized. Details kept non-public.
August 4, 2026
Staff-level briefing with major AI labs. Framework not released publicly. Benchmarks classified.
What We Know (and What We Do Not)
The framework itself is unclassified, but the specific thresholds and benchmarks are not public. Here is what has been confirmed by reporting from Fortune, Axios, Next Web, and CNBC, and what remains opaque.
Confirmed
- +A 30-day pre-release window applies to frontier models with advanced cybersecurity capabilities
- +The NSA plays a central role in model capability assessments
- +OpenAI, Google, and Anthropic reviewed a draft and submitted edits
- +Open-weight models are explicitly excluded from the framework
- +The framework is voluntary: it does not carry regulatory enforcement authority at launch
- +The framework targets models with advanced capabilities in cybersecurity and potentially CBRN domains
Not publicly known
- ?Specific capability thresholds that trigger the pre-release review requirement
- ?Which benchmarks the government uses to evaluate frontier capability
- ?What happens procedurally during the 30-day window
- ?Whether the review produces a pass/fail outcome or a conditional release
- ?Whether smaller labs not in the August 4 briefing are covered by the same framework
- ?The timeline for the framework becoming mandatory rather than voluntary
Who This Actually Affects Right Now
The most important thing most AI PMs can do with this news is correctly identify whether it applies to their work. The current framework is narrow in scope. Most AI product teams do not build frontier models and are not subject to the pre-release review requirement.
Directly affected: frontier model labs
OpenAI, Anthropic, Google, Meta, xAI, and any other organization training models at or near frontier capability in cybersecurity domains. These teams need to build the pre-release review process into their release roadmaps. A 30-day government review window is a material constraint on release cadence.
Potentially affected: advanced security AI teams
Teams building AI-powered cybersecurity tools that approach frontier offensive capability: automated vulnerability discovery, AI-assisted exploit development, advanced penetration testing automation. The framework's cybersecurity focus suggests these products may eventually come under review as their models scale.
Not affected: application-layer AI teams
Teams building AI products on top of API-accessible frontier models (ChatGPT, Claude, Gemini) are not directly covered. The review requirement targets the models themselves, not the applications built on them. If your product calls an OpenAI API, the responsibility lies with OpenAI, not with your application team.
Explicitly excluded: open-weight model projects
Open-weight models (Llama 4, Qwen3.8, Muse Glimmer) are explicitly outside the framework's scope. The administration has acknowledged it cannot enforce pre-release review on models whose weights are publicly released: once weights are public, the review window is irrelevant.
Navigate AI Regulation Without Slowing Down
The AI PM Masterclass covers how to build compliance requirements into your product roadmap without losing velocity, taught by a former Apple and Salesforce Sr. Director PM.
Strategic Risks of a Secret Standard
Keeping the evaluation framework secret solves one problem (adversarial calibration to just below the review threshold) and creates several others. AI PMs at model labs need to reason about these strategic risks in their release planning.
You cannot benchmark against an unknown standard
If you do not know which capability thresholds trigger the review window, you cannot self-assess whether your model requires government review before release. Labs may either over-report (submit models that do not need review, adding unnecessary delay) or under-report (miss the trigger and face regulatory consequences post-release).
The voluntary status is temporary
The current framework is explicitly voluntary. It does not carry enforcement authority. But voluntary frameworks with government backing have a consistent historical trajectory: they become mandatory, often rapidly. Build the pre-release review process into your release workflow now, before it is required.
US allies outside the briefing are disadvantaged
AI labs in allied nations that were not included in the August 4 briefing are making release decisions without access to the same framework that US labs reviewed. This creates a compliance asymmetry: US labs have informal guidance; international labs do not. Expect diplomatic pressure to change this.
Open-weight exclusion creates a structural gap
The framework explicitly excludes open-weight models. If a foreign adversary open-sources a frontier model with advanced cybersecurity capabilities, the framework provides no mechanism to respond. This gap is already a subject of internal government debate and is likely to produce framework revisions.
NSA involvement changes the compliance calculus
Most product and legal teams have extensive experience with FDA, FTC, and SEC compliance frameworks. NSA-led technical assessments are structurally different: the standards are classified, the reviewers are intelligence community professionals, and the process is not subject to the same transparency requirements as civilian regulatory review.
What Product Teams Should Do Now
The right response to this development depends entirely on which category your team falls into. Here are concrete actions for each group.
If you work at a frontier model lab
- 1.Add a 30-day pre-release review buffer to every major model release roadmap for any model with advanced cybersecurity capability
- 2.Assign a government affairs lead to maintain active communication with the OSTP and relevant agencies about where your models sit relative to the (unknown) thresholds
- 3.Begin building internal capability assessment documentation now: behavioral logs, red team results, safety evaluations. If the government requests documentation during a 30-day window, you want it pre-prepared
- 4.Engage legal counsel with experience in national security compliance frameworks to advise on the NSA review process
If you build security-adjacent AI products
- 1.Monitor the framework's evolution. As the framework matures, the scope is likely to expand. Track OSTP communications and Congressional hearings on AI regulation for early signals
- 2.Document the capability limits of your AI features explicitly. If a regulator asks whether your product has frontier cybersecurity capability, you want a pre-built answer backed by evidence
- 3.Avoid building capabilities that could be interpreted as crossing into frontier offensive cybersecurity territory without a clear legal and product rationale
If you build application-layer AI products
- 1.No immediate action required on this specific framework
- 2.Understand that your API providers (OpenAI, Anthropic, Google) are navigating this. Model updates from these providers may become slower if government review adds time to their release cycle
- 3.Monitor your providers' public statements about the framework: they will communicate material changes to their release cadence if the review window affects their timeline
The Bigger Picture: AI Governance Is Entering a New Phase
The White House's secret AI evaluation framework is a symptom of a broader shift in how governments relate to frontier AI development. For the past several years, AI governance was primarily a standards-setting exercise: agencies and legislative bodies defining principles, publishing guidelines, and hoping industry would self-regulate. That era is ending.
Parallel development: EU and US are converging
The EU AI Act's high-risk provisions activated August 2, one day before the US framework was finalized. Two of the world's largest regulatory jurisdictions are now running mandatory and semi-mandatory AI review processes simultaneously. Any global AI product faces compliance obligations in both regimes.
NSA involvement normalizes intelligence community oversight
Including NSA as a central assessor in a civilian AI evaluation framework is structurally unprecedented. It signals that the government views frontier AI capability primarily through a national security lens, not a consumer protection or competition lens. Expect this framing to shape future legislation.
The voluntary-to-mandatory pipeline is accelerating
The gap between voluntary frameworks and mandatory regulations has historically been 2 to 5 years for emerging technology sectors. For AI, given the pace of capability development, that gap is compressing. Product roadmaps that plan for mandatory pre-release review within 18 months are not being paranoid.
Regulatory differentiation is a product strategy
Companies that build the capability to comply with pre-release review requirements before they are mandatory gain a structural advantage when compliance becomes required: they have the processes, documentation, and government relationships that competitors lack. Compliance infrastructure is a moat.
Build AI Products That Thrive Through Regulatory Change
The AI PM Masterclass covers how to monitor regulatory developments, embed compliance requirements into product roadmaps, and turn governance constraints into competitive advantages.
Related Articles
Before you go: get the AI PM Minute
One tactic to make you a sharper AI PM, twice a week. 60 seconds to read. Free.
No fluff. Unsubscribe anytime.