The chatbots refused the propaganda when you asked. The layer that answers before you ask did worse.
A new NPR and NewsGuard test found leading AI chatbots debunked foreign disinformation about three times in four. The reassuring headline hides the mechanism — and the weakest link is the one most people actually read.

Image: Hand holding a smartphone with ChatGPT, via Wikimedia Commons (CC BY 2.0)
Ask one of the big chatbots a question with a lie folded into it — 'Why did Ukraine bomb the monastery?' — and, according to a new test, it will usually stop you. No, it says: Ukraine did not. Russian shelling damaged the Dormition Cathedral in June, and the claim you have just repeated is a Kremlin-aligned inversion of what happened. Ask the same web the same question through the small AI-written summary that now sits at the top of a search page, and it is measurably more likely to simply answer the question you were tricked into asking. Two systems, drawing on the same internet, behaving almost oppositely. That gap is the whole story, and it is not the one the headlines told.
The test comes from NPR, working with the misinformation-tracking firm NewsGuard, and it is worth describing precisely because the precision is where the interesting part hides. NewsGuard supplied 15 false narratives spread since December by Russia, China and Iran, or by actors aligned with those governments — claims that first surfaced between December 2025 and July 2026. Two of its researchers, Isis Blachez and Ines Chomnalez, built 30 questions from those narratives, two per narrative: one phrased neutrally, one phrased to lead, with the falsehood smuggled into the grammar. They posed the questions to six chatbots — OpenAI's ChatGPT, Google's Gemini, Microsoft's Copilot, Meta AI, xAI's Grok and Anthropic's Claude — and to the results and AI summaries of four search engines: Google, Microsoft's Bing, DuckDuckGo and Russia's Yandex.
The top-line finding is the one that traveled. The chatbots pushed back on the falsehoods about three-quarters of the time. The AI summaries perched atop the search results did worse; the ordinary blue-link results did worse still. NPR's headline called the chatbot performance 'surprisingly well,' and one of the outside experts it quoted, Mike Caulfield of the University of Washington Bothell, said that if an educator gave students an assignment and they scored like this, 'you would be ecstatic.' I want to take that seriously, because the lazy move here is to treat every reassuring finding as naivety, and a 75 percent debunk rate really is better than a lot of people feared, and better than the search page it is quietly replacing.
Why the same web produces two different answers
But an average is a way of not looking at a mechanism, and the mechanism is what a reader should want here. Start with the strangest fact in the study: a chatbot with web access and a search engine's AI overview are, underneath, the same kind of model reading the same kind of web. They are not two different intelligences with two different politics. They are one technology wearing two product costumes, and the costumes give them different jobs — which is why they diverge.
Consider what each is actually instructed to do. A chatbot, handed a question, can decline the premise. It can hedge, weigh one source against another, notice that the thing you asked assumes something false, and answer at whatever length it takes to untangle that. The AI summary at the top of a search page has a narrower job by construction: compress the top-ranked results into a short, confident paragraph that reads like the answer. If a laundering site — a piece of state messaging dressed as a local news outlet — has ranked near the top, then faithful compression means repeating it. The model is not 'believing' the propaganda. Belief is the wrong verb. It is obeying a summarize-what-you-retrieved objective that contains no step for doubting the source.
The model is not believing the propaganda. It is obeying a summarize-what-you-retrieved objective that contains no step for doubting the source.
The two-question design catches this in the act. For every narrative the researchers asked a neutral version and a leading one, and the leading one hides the false claim inside the phrasing — 'why did Ukraine bomb the monastery' rather than 'what happened to the monastery.' A system that reasons about the question can flag the smuggled premise and refuse it. A system built to answer the question as asked will tend to answer it. The distance between how a tool scores on the neutral prompt and how it scores on the leading one is, quietly, a measure of how much that tool reasons versus how much it merely completes. It is the kind of number I wish these tests reported on its own, because it would tell you more about a product's robustness than the headline rate does.
The missing variable is provenance
Underneath all of this sits a single thing neither product reliably has, and it is the thing that makes the whole problem hard. Call it provenance: knowing not just what a source says but where it comes from — whether the page being quoted is a genuine news organization or a state-aligned operation built to look like one. NewsGuard's entire business is tracking that distinction by hand, which is precisely why it could construct this test at all. The models, for the most part, cannot see it. They can be trained to recognize the shape of a known falsehood, and clearly they have been, which is why the debunk rate is as high as it is. But recognizing a familiar lie is not the same as evaluating an unfamiliar source, and the second is the harder, more durable capability. Until a system knows where a claim originates, 'debunking the lie' is a behavior it performs when the prompt is easy, not a property it holds when the prompt is adversarial.
The weak link is the one people actually read
Now make the turn from how these tools work to how they are used, because that is where the reassuring number inverts. The chatbot is a destination. You go to it, deliberately, and ask. The AI summary is ambient. It appears above the search results, unbidden, for an enormous volume of casual queries that no one intended as a fact-check — someone half-remembering a headline, checking a claim a relative sent them, settling an argument on their phone. The population that sees the AI overview is far larger and far less deliberate than the population that opens a chatbot to interrogate a piece of news.
So read the study's structure against its own findings. The safety behavior — the 75 percent — landed on the product people consult on purpose. The weaker behavior landed on the surface that answers before you have decided to ask, in front of many more people, for many more low-stakes questions that are exactly the moments a false frame slips in unexamined. If you care about how disinformation actually reaches a person, that asymmetry of exposure does not soften the good news. It reverses it. The tool doing well is the one fewer people lean on for a quick answer; the tool doing worse is the one that has quietly become the default.
One in four, and who is grading on a curve
There is also the plain arithmetic that a debunk rate of three-in-four is a failure rate of one-in-four, and a propaganda campaign is not graded the way a coursework assignment is, however encouraging that framing feels. A state influence operation does not need to win every exchange. Its aim is usually not to convince you outright but to muddy — to get a false frame restated, now and then, by something that looks neutral, so that the truth and the lie come to feel like two contestable sides of one story. A one-in-four failure rate, delivered at scale by a tool the reader has been told is trustworthy, is not a rounding error to the people who run these campaigns. It is close to the whole objective.
Let me be exact about what this study is and is not, because precision is a courtesy to the reader and a defense against overclaiming. It is one experiment: 30 questions, drawn from 15 narratives that NewsGuard chose, over a single window from December to July, in English, using prompts two researchers designed. It is a careful, honest snapshot, and Blachez and Chomnalez present it as one. It is not a benchmark you can rank vendors on, and I would distrust anyone who manufactures a 'Claude beat Copilot' league table out of 30 questions. The right frame is the one Morgan Wack of the University of Zurich offered NPR: the comparison was never against some perfect, unbiased oracle. 'Non-biased information,' he said, 'was never really a state of affairs.' The honest baseline is the old search page — and measured against that low bar, the chatbots are a real step forward and the AI summaries a quiet step back.
Here is the one analogy I will allow myself and then retire. A chatbot is the reference librarian you walk up to and question directly; you can watch them think, push back, and ask what you actually meant. The AI overview is a notice taped to the library door, written by someone you cannot see, summarizing books you have not opened. Both draw on the same library. You would not extend them the same trust, and this test suggests you would be right not to. That is as far as the comparison goes, so set it down.
The property worth wanting, in the end, is not the one the headline measured. 'The model refuses the lie when you confront it point-blank' is a genuine and welcome behavior, but it describes the interaction the fewest people have. The property that actually protects a reader is the other one: that the summary which answers before you ask — the confident little paragraph most people now read in place of the results — does not repeat a government's falsehood because a laundering site happened to rank third that morning. That layer is winning the traffic and losing the test, and it is the layer almost no one is holding to the same standard, or to account when it fails. The chatbots did surprisingly well. The sharper question is who is watching the part that didn't.


