Opinion · AI governance

Washington found a way to review AI models for national security without licensing them. That is the problem, not the fix.

A voluntary framework with OpenAI, Anthropic and Google, expected within days, lets federal evaluators test frontier models before release. It is neither the licensing regime its critics fear nor the voluntary standard its authors claim — and that third thing answers to no one.

The west front of the United States Capitol under a clear sky.

Image: Architect of the Capitol / Wikimedia Commons (public domain)

Within the next week or so, the White House is expected to finalize a voluntary framework under which OpenAI, Anthropic and Google — along with Microsoft and Amazon — will give federal evaluators up to thirty days with each new frontier model before the public gets it, so the government can test what the model can do in a cyberattack. Two camps have already decided what this is. To one, it is what safety advocates have wanted for years: at last, someone checks these systems before they ship. To the other, it is the soft opening of a licensing regime — an FDA for software, preclearance by another name. Both are wrong, and they are wrong in the same way. I say this as someone who spent most of a decade inside agencies writing rules of roughly this kind, and who, in another life, would have been in the room drafting this one.

The category error is the argument itself. Everyone is asking whether this is a license or whether it is voluntary, and the honest answer is that it is neither, which is precisely what makes it worth worrying about. Start with the text, because the drafters were careful with it. The executive order that set this in motion, signed on the second of June, explicitly prohibits the government from establishing a ‘mandatory governmental licensing, preclearance, or permitting requirement’ for developing or releasing a model. Take them at their word: they went out of their way not to build a license. There is no approval you must obtain, no permit, no legal gate you cannot pass without a stamp.

But it is not a voluntary standard either, and pretending otherwise is the tell. Participation runs on leverage — the same export-control authority that, over the past two months, has been used to restrict who may use these companies' most capable models. When a firm knows the government controls whether its product can reach foreign markets, an invitation to submit that product for review is not really an invitation. One analyst put it more plainly than the officials will: voluntary participation in a framework backed by export-control authority is participation under pressure, whatever the order's text says. So this is the third thing, the one our vocabulary does not have a clean word for: a classified, discretionary national-security review that binds through leverage rather than law. A thing without a name escapes the accountability that names carry, and that is not an accident of drafting. It is the design.

The strongest case for it, stated fairly

Let me give the government's case its best form, because it deserves one. Some capabilities really are national-security matters and not consumer-safety matters. A model that can autonomously discover and weaponize software vulnerabilities at machine speed is not a defective toaster; it is closer to a munition, and we do not run munitions through notice-and-comment. The people who can actually assess that capability — at the National Security Agency, and at the new Center for AI Standards and Innovation inside the Commerce Department — have genuine, hard-won expertise that no public rulemaking docket could assemble, and some of what they know cannot be published without teaching the exact thing they are trying to contain. Thirty days is modest; it is a look, not a veto, and the recent record shows the mechanism operating with some restraint — the newest models went out to small sets of vetted partners during review rather than being blocked outright. If you believe, as I partly do, that frontier cyber capability is a real and near-term danger, a brief expert look before wide release is not a crazy thing to want. That is the steelman, and it is strong enough that the reflexive libertarian answer — get out of the way — does not survive contact with it.

The problem is not the wanting. It is the machine that has been built to do it, which fails on exactly the procedural grounds that decide whether a rule is legitimate or merely powerful.

Three ways the machine fails the test

First, the benchmark is classified. This is the load-bearing objection and everything else follows from it. The labs will not know precisely what their systems are being tested against; independent researchers cannot audit the results; no court can review a finding it is not allowed to see. Compare that to how a real rule works, even a tough and technical one: the agency publishes the standard, so the regulated party can conform its conduct to it and a judge can check whether the agency actually applied its own test. Strip out the published standard and ‘review’ stops being the application of a rule and becomes the exercise of a discretion — whatever the reviewer decides this quarter, unfalsifiable from the outside. You cannot conform to a test you cannot read, and you cannot challenge a judgment you cannot see. That is not a standard. It is a mood with an infrastructure.

You cannot conform to a test you cannot read, and you cannot challenge a judgment you cannot see. That is not a standard. It is a mood with an infrastructure.

Second, who is in the room is the process. According to the reporting, Commerce Secretary Howard Lutnick personally approved roughly twenty customers during a limited preview of one company's model. Sit with that detail, because it is the whole thing in miniature. Not an office applying a published criterion — a cabinet secretary, by hand, deciding who may buy access to a commercial product. That is not enforcement; it is patronage. When the decision turns on the identity and mood of the official rather than on a written rule, the rule is the official, and it changes when he does. The entire discipline of administrative law — the unfashionable, deeply practical field of how agencies are supposed to make decisions — exists to prevent precisely this: to force the outcome to rest on a criterion anyone can read rather than on whose hand is holding the pen. This framework dispenses with that discipline and calls the result agility.

Third, it binds only the willing, which means it is calibrated to cooperation rather than to risk. Meta is not in the deal, because Meta releases its strongest models as open weights, and you cannot give a government thirty days’ head start on a model that anyone can already download. That is the tell, and it is fatal to the framework's own theory. The capability the review most fears — a frontier model that can be copied, fine-tuned to strip its safeguards, and re-hosted beyond anyone's reach — is exactly the one the framework cannot touch, because it lives outside the set of firms that agreed to be reviewed. So the review lands hardest on the most cooperative labs and not at all on the distribution model that keeps the people who wrote it up at night. A safety regime shaped like that is not calibrated to danger. It is calibrated to who showed up to the meeting, and the most dangerous actor, by the framework's own logic, is the one who stayed home.

Who enforces this, and how

Ask the question I always end up asking, the one the slogans skip: who actually enforces this, and how? A genuinely voluntary standard is enforced by nothing but reputation. This one is enforced by the leverage that makes it non-voluntary in the first place — export controls, delayed approvals, the quiet phone call — instruments that are discretionary, largely shielded from review, and applied case by case behind a classification. The enforcement mechanism and the accountability gap are the same object. A rule whose only teeth are invisible pressure cannot be checked when the pressure is misapplied, because from the outside you cannot tell whether it was applied at all. You will not see the model that was quietly discouraged, the release that slipped for reasons no one will state, the competitor who did not get Secretary Lutnick's twenty slots. The absence of a paper trail is not a side effect here. It is the product.

The better frame, and what it costs

If the cyber risk is real enough to gate frontier models — and I am prepared to believe it is — then it is real enough to govern in the open, and the fact that open governance is harder is not a reason to prefer the version that answers to no one. Concretely, that means four things. A capability threshold published in advance, so a lab knows what triggers review instead of guessing. A test whose methodology is public even where specific attack payloads stay classified — we manage exactly this split in other national-security domains, where the existence and shape of a control are public and only its contents are secret. A statutory basis, so the authority belongs to Congress and survives the administration that created it. And judicial review, even in a cleared and closed setting, so that a finding can be contested by the party it lands on. Above all, a rule that attaches to a capability rather than to the five companies that came to the table — which means confronting the open-weights problem honestly instead of quietly routing around it and hoping no one notices the hole.

I will not pretend that path is costless, because the discipline of this column is to admit what my own position gives up. A published threshold can be gamed right to its edge. A statute takes a year or more that the technology will not politely wait through. Doing this in the open might mean a genuinely dangerous model reaches the world during the gap while the slower machinery catches up. Those are real risks and I am not going to wave them away. But weigh them against what the current design actually builds: a permanent, classified, discretionary national-security review over software, run by whoever holds the office, binding only the compliant, with no published rule, no audit, and no appeal. The first path risks a bad model slipping out once. The second normalizes a bad machine for good — and machines, unlike models, do not get recalled.

That is the choice underneath the one being debated, and it is why I cannot join either camp cheering or damning ‘the AI licensing regime.’ It is not a licensing regime; it is something with less accountability than a license and more reach. The thing about a review built to be unreviewable is that it works exactly as well for a bad administration as for a good one, and the people building it right now do not get to choose which one inherits it. Neither, when the time comes, will the rest of us.

References

  1. Eastern Herald: White House and top AI labs near deal on voluntary frontier-model standards
  2. The Hill: Microsoft, Google, xAI giving government early access to AI models for review
  3. TechTimes: Trump AI order creates voluntary 30-day review window for frontier models
  4. SecurityWeek: OpenAI and Anthropic limit new AI models to Trump-approved customers during cybersecurity review
  5. AI Weekly: White House nears voluntary frontier-model deal with top AI labs
The Friday Brief

One email. Every Friday.

The week's machines, money, and people — in under five minutes.