Washington just built a test that decides which AI models are too capable to ship freely. The passing grade is classified.
An executive order's August 1 deadline put a secret National Security Agency benchmark at the center of frontier AI. It defines a category no company can read, run by an agency built for signals intelligence — and the order calls it voluntary.

Image: National Security Agency (public domain)
The most consequential document in American AI policy this summer is one almost no one is cleared to read. It is a benchmark — a test that sorts artificial-intelligence models by how dangerous their capabilities are judged to be — and as of August 1 it sits at the center of how the most advanced models in the country reach the public. It was due that day under an executive order signed two months earlier. Its contents are classified. So is the threshold it sets. A company cannot read the standard its own model will be measured against, and neither can you.
The order is Executive Order 14409, "Promoting Advanced Artificial Intelligence Innovation and Security," signed June 2, 2026. It gave federal agencies sixty days — a clock that ran out August 1 — to stand up two things. The first is "a classified benchmarking process for assessing the advanced cyber capabilities of AI models," according to the order, used to determine which systems count as "covered frontier models." The second is a voluntary framework under which the developer of a covered model may give the government access for "up to 30 days before the developer's planned release date" to "other trusted partners." Read those two provisions together and a structure appears that the order's own language works hard to keep from naming.
What the order says on paper
On its face, EO 14409 is a light-touch document, and its drafters say so. The order "expressly states that it does not create any mandatory governmental licensing, pre-clearance, or permitting requirement," in the summary of one law firm that advises AI developers. Participation in the pre-release review is voluntary. There is no statute here compelling a company to hand over a model, and no license a company must obtain before shipping one. In the administration's telling, this is the market-friendly alternative to the licensing regimes floated in Washington and enacted in Brussels: a benchmark, an invitation, and a thirty-day window a developer may choose to open.
The word carrying the weight is voluntary, and it is worth holding up to the light. A standard is voluntary until it becomes the only accepted way to do a thing. The order contemplates that covered frontier models will, after government review, be released to "trusted partners" — a category the order does not define, but which its drafters read to include the critical-infrastructure operators that run power, water, finance and health systems. If clearing the classified benchmark becomes the path to selling advanced AI into those buyers, then "voluntary" describes the paperwork, not the incentive. One trade-press analysis put the tension in its headline: voluntary on paper, mandatory in practice.
A company cannot read the standard its own model will be measured against. The category that triggers the review is itself classified.
The definition is the mechanism
Here is the part that does the real work, and it is easy to miss because it hides inside a definition. The order does not say what a "covered frontier model" is. It does not set a parameter count, a compute threshold, or a capability line. Instead, it assigns three departments — the Treasury, the Department of War through the National Security Agency, and the Department of Homeland Security through its Cybersecurity and Infrastructure Security Agency — to establish that threshold through the classified benchmarking process itself. The test defines the category. And the test is secret.
Follow what that does to a developer. To know whether your model is "covered" — whether the thirty-day pre-release review applies to you at all — you must know whether it crosses a line drawn inside a classified process you are not cleared to see. The trigger and the standard are the same undisclosed object. In ordinary regulation, the definition of the regulated thing is public even when the enforcement is tough; you can read the statute that names you. Here the naming happens behind a classification stamp. The most basic question in any rulebook — does this apply to me? — is answered by a document its subjects cannot open.
Who holds the pen
The order distributes the work, but not evenly. The Treasury is tasked with an AI cybersecurity clearinghouse. Homeland Security, through CISA, is charged with building the pre-release framework. But authority over the classified benchmark — the piece that sets the threshold — runs through the Secretary of War to the National Security Agency. The Center for AI Standards and Innovation, the federal body formerly known as the AI Safety Institute, is named among the reviewers. Drafts of the framework were reportedly circulated to OpenAI, Anthropic and Google as it was written.
The NSA's role is the detail worth sitting with, because of what the agency is. The NSA is a signals-intelligence organization. Its mission is to break into communications systems and to defend the government's own — not to weigh consumer harm, competition, or civil liberties, which are the concerns that usually attend a decision about what a company may sell. Legal analysts tracking the order flagged the same point: the NSA's position as the primary decision-maker over so consequential a process, one commentary noted, "stands out." A benchmark for "advanced cyber capabilities" is squarely inside the agency's expertise. Whether that expertise should decide which commercial AI models reach the market is a different question, and the order answers it by default rather than by argument.
The thirty days, and the asymmetry
Then there is the window. Under the framework, the government may examine a covered frontier model for up to thirty days before it goes to trusted partners. Set aside whether thirty days is long or short. Look at the direction of the flow. The government — through an intelligence agency — gets to see and probe a frontier model before the companies that will deploy it, before the researchers who study these systems, before the public that will use them. What flows the other way is far thinner. The public does not get to see the model early. It does not get to see the benchmark at all. It does not get to see the threshold, the methodology, the score, or the reasoning behind a finding that a given system is or is not "covered."
This is the asymmetry that runs through every surveillance-adjacent system I have covered, arriving here in a new setting: one side accumulates visibility, the other accumulates exposure, and the arrangement is described in language calm enough to make the imbalance sound procedural. The government will know things about these models that their own makers' customers will not. The standard by which it judges them is not merely undisclosed; it is classified, which means it is unreviewable by design — not withheld pending release, but built to stay dark.
What a classified standard forecloses
A classified benchmark cannot be audited from outside. That is not a criticism of its authors' competence; it is a property of the form. There is no published methodology to contest, no threshold to litigate, no score to appeal, no paper trail a journalist or a court can request. If a model is judged too capable to ship freely, there is no public record of why. If a model that should have been caught is waved through, there is no public record of that either. The category of "covered frontier model" will grow or shrink according to a document that leaves no trace the governed can follow. Accountability runs on records. This process is designed to keep the decisive ones sealed.
OpenAI's forthcoming Astra system — the model family the company began previewing to policymakers in Washington this week — is expected to be among the first submitted under the framework. It will be a useful test of how the process behaves in daylight, which is to say, of how little of it is visible in daylight at all. The company will disclose what it chooses. The government will disclose what it must, which under a classification stamp is close to nothing.
None of this requires the tradecraft inside an NSA cyber benchmark to be published; some of it genuinely cannot be, and pretending otherwise would be its own kind of dishonesty. But there is a wide gap between the secrets that protect a method and the secrets that hide a decision. The order could name, in public, who sets the threshold and on what statutory authority; the conditions under which a model becomes "covered"; who "trusted partners" are; and what recourse a developer has when the finding goes against it. It publishes none of these. The result, as of August 1, is a regime whose central instrument is a test the regulated cannot read, defining a category they cannot check, administered by an agency whose business is secrets. The government has given itself thirty days to look inside the most capable models being built in America. It has given everyone else a stamp that says classified, and asked them to trust that the line is being drawn somewhere fair.
References
- Latham & Watkins — President Trump Signs Executive Order Establishing AI Cybersecurity and Frontier Model Framework
- Wiley — New AI Executive Order Addresses Frontier Models and Cybersecurity Vulnerabilities
- Congress.gov (CRS) — Controlling Advanced Artificial Intelligence: Executive Order 14409 Explained
- TechTimes — Voluntary on Paper, Mandatory in Practice: White House AI Review Hits August 1 Deadline
- Cato Institute — Assessing the Trump Administration's Executive Order on AI and Cybersecurity


