Surveillance

The Pentagon put three million people's sensitive work inside commercial AI. The safeguard is a sentence in the terms.

GenAI.mil now routes controlled unclassified work through ChatGPT, Grok and Gemini. What protects that data is not a wall — it is an accreditation boundary and a promise not to train on it.

Aerial view of the Pentagon, headquarters of the U.S. Department of Defense

Image: U.S. Navy photo (030926-F-2828D-089) / Wikimedia Commons, public domain

On Monday the Pentagon added two commercial chatbots to a portal that up to three million military and civilian personnel are meant to use, and the most important sentence in the announcement was not about capability. It was about custody. ‘Data processed on GenAI.mil remains isolated within the government environment and is not used to train or improve its public or commercial models.' That is the safeguard. Not a wall, not an air gap. A sentence, in a set of terms, describing what a vendor has agreed not to do with information it is now positioned to see.

The portal is GenAI.mil, the Defense Department's centralized front door to commercial frontier AI. It launched roughly nine months ago with Google Gemini. As of Monday it also offers ChatGPT Mil, from OpenAI, and Grok for Government, from Elon Musk's xAI and its Starshield defense arm. The Defense Department says the point of running three models side by side is to ‘eliminate vendor lock-in and promote a broader American AI ecosystem.' More than 1.7 million people have used the platform since it opened. The stated ceiling is the full workforce: around three million.

This is not a story about whether soldiers should use AI. They already are. It is a story about where a specific, enormous body of sensitive-but-unclassified government information now travels, who is contractually allowed to touch it, and what exactly stands between it and the training pipelines of the largest AI companies in the country. Read the accreditation and the guidance, link by link, and the answer is more fragile than the announcement suggests.

What the accreditation actually covers

Start with the boundary the government drew. ChatGPT Mil is accredited to handle Controlled Unclassified Information at Impact Level 5, a Defense Department cloud-security standard for certain sensitive, non-public data. Grok for Government cleared the same bar. That accreditation is the whole permission structure. It is what lets a service member paste something into a chatbot that they could not paste into the consumer version of the same product. IL5 is a real and demanding standard. It is also a scope, and a scope is defined as much by what it excludes as by what it covers.

Controlled Unclassified Information is a vast category. It is logistics and planning and policy drafts and acquisition analysis and supply-chain reconciliation — exactly the document-heavy administrative work the Defense Department says the tools are for. It is not classified material, and nothing here suggests classified material is going anywhere near these systems. But CUI is precisely the tier of information that is sensitive enough to hurt and mundane enough to be handled all day by people who are busy. The accreditation says this data may flow through a commercial model. It does not, by itself, say what the model may remember, what its operator may log, or how long any of it persists. Those are separate questions, answered in separate documents, by the vendors.

IL5 is a real standard. It is also a scope. And a scope is defined as much by what it excludes as by what it covers.

The promise not to train is a contract, not a mechanism

Here is the load-bearing claim, in the words of OpenAI's own government sales chief, Joe Larson: data processed on the platform ‘remains isolated within the government environment and is not used to train or improve' the company's public or commercial models. Take that at face value. It is a meaningful commitment, and there is no evidence on the record that any vendor intends to break it. But notice what kind of protection it is. It is a promise. It is enforced by contract terms and configuration, not by physics. The difference matters, because the entire history of the data economy is the history of promises like this one being technically true and practically porous — data ‘isolated' in an environment the vendor still operates, ‘not used for training' under a definition of training the vendor writes.

The reporting is careful, and so is the language: military versions are exempted from the ordinary consumer data-collection practices that apply when you or I use these products. Exempted. That word is the tell. It means the default behavior of these systems is to collect, and the government version is a negotiated departure from that default — a departure that holds exactly as long as the configuration holds and the contract is honored. When the protection is a carve-out from a data-hungry default, the thing to watch is not the promise. It is the machinery underneath it, which was built to do the opposite.

There is a smaller document that says more than the press release. Navy guidance already instructs personnel not to put personally identifiable information or protected health information into the platform at all. Read that plainly: the institution deploying these tools is telling its own people not to trust them with the most sensitive personal categories. That is not a criticism buried by a critic. It is the operator's own instruction, and it is an admission that the safeguard for the highest-sensitivity data is not the accreditation or the contract. It is a rule asking three million busy people not to type the wrong thing — the least reliable control there is.

Who got in, and who was kept out

Follow the roster, because the roster is an argument. Inside the tent now: OpenAI, Google, and xAI, whose Grok is positioned for everything from acquisition market research to supply-chain management. Around it, the Defense Department has signed separate AI arrangements with Amazon Web Services, Microsoft, Nvidia, and Reflection AI. That is most of the American frontier-AI industry, wired into the workflow of the largest employer in the country.

One major lab is conspicuously absent: Anthropic, maker of Claude. The reason is a matter of record. The Trump administration designated Anthropic a supply-chain risk after the company refused to grant the Pentagon unrestricted access and insisted on keeping certain safety guardrails in place; Anthropic is challenging that designation in court. Set the litigation aside and look only at the pattern the two facts make together. The vendor that pushed back on how its model could be used is the one locked out. The vendors now handling millions of people's sensitive work are the ones that cleared the bar the government set. Whatever else that is, it is a selection pressure, and it selects for accommodation.

It is worth being precise about the one name that draws the most heat. Grok, from xAI, has a public track record of safety failures on its consumer product that has been documented repeatedly over the past year. Its government version is a different, accredited deployment, and the consumer incidents do not automatically transfer to it. But the burden of that distinction runs the other way from how it is usually told. The government is not inheriting the consumer Grok's problems; it is trusting that a separately configured version does not carry them, on the same basis it trusts everything else here — the vendor's assurance and the accreditation, not an independent look under the hood that the public can see.

The structure, stated plainly

Assemble the pieces and the shape is clear. The Defense Department set out to solve a real problem — dependence on any single AI vendor — and solved it by multiplying the number of commercial models sitting between a service member and their own sensitive work. Vendor lock-in goes down. Vendor exposure goes up. Where before there was one company's terms to trust, there are now several, each with its own configuration, its own logging, its own definition of what ‘not used for training' means in practice.

None of this requires anyone to be acting in bad faith. That is the point. Every link in the chain is individually defensible: an accreditation that genuinely raises the bar, a contract clause that genuinely limits use, a guidance memo that genuinely warns people off the riskiest data. Collectively they add up to a system in which the sensitive administrative life of the world's largest military runs through commercial AI, and the thing protecting it is a stack of promises and a rule against typing the wrong thing. The data economy has always run on the assumption that no one will assemble the whole picture. Here the whole picture is not hidden. It is in the accreditation memos and the vendor statements, sitting in public, waiting to be read as a system instead of a launch.

What would change the calculus is not more models. It is verification the public and Congress can see: independent audits of whether ‘isolated' data stays isolated, published definitions of what each vendor counts as training, retention limits with teeth, and an accounting of who can access the logs. Until then, the safeguard for three million people's sensitive work is what it was on Monday — a sentence in the terms, and the good faith of the companies that wrote it.

References

  1. Military Times — The military's ChatGPT is now live via the Pentagon's GenAI platform
  2. DefenseScoop — Grok, ChatGPT added to GenAI.mil
  3. TechCrunch — The Pentagon now has its own version of ChatGPT and Grok
  4. The Hill — Pentagon adds Grok for Government and ChatGPT for military use
  5. Hoodline — Pentagon adds ChatGPT and Grok to GenAI.mil, reaching 3 million users
The Friday Brief

One email. Every Friday.

The week's machines, money, and people — in under five minutes.