The model that answers every Google search is the cheapest one Google has
Google now writes an AI answer to every query, and the model doing it — Gemini Flash — is tuned for speed and cost, not for being checked. When one system mediates the world's questions, the thing that disappears is provenance.

Image: The Pancake of Heaven! / Wikimedia Commons (CC BY-SA 4.0)
When you ask Google a question now, a language model writes the answer. Not a ranked list of pages that might contain it — the answer itself, composed on demand, in the moment you ask. Since the tenth of July, by the company's own account, this is no longer a feature you switch on but the default behaviour of the search box, applied across Google's surface, worldwide. The interesting fact is not that a model writes the answer. It is which model. It is not the largest or most capable system Google has built. It is close to the cheapest one that can keep up.
The model doing the writing is Gemini Flash — the fast, inexpensive tier of Google's line, built to return a lot of tokens quickly at low cost per token. This week Google shipped its latest version, Gemini 3.6 Flash, and the detail worth pausing on is how the company described the improvement: the new model produces, by its own figure, seventeen percent fewer output tokens than the one before it, at a lower price. That is the direction of travel stated out loud. The model that answers the world's questions is being tuned, generation over generation, to say a little less for a little less money. Flash is not the model Google would reach for if the only goal were to be right. It is the model you reach for when you have to answer billions of queries a day and the arithmetic of doing that with anything heavier does not close.
It is worth being precise about why the heavier model is not the one on the search box, because the reason is not sinister; it is structural. Google's flagship, Gemini 3.5 Pro, was reported by Bloomberg this month to have been delayed after missing internal targets. But even had it shipped on time, it was never going to be the thing answering every query. A frontier model is expensive to run per token, and search volume is the largest sustained inference load any company has ever faced. The model that mediates it has to be selected first for speed and cost, and only then for everything else. That ordering is the whole story. It is not that Google chose a worse model out of carelessness. It is that at planetary scale, cost is not one consideration among many. It is the constraint that picks the model, and the model it picks is the one optimised to be brief and cheap.
What a generated answer removes
To see what changes, hold the two things side by side — not the old world and the new, but what each actually hands you when you ask it something.
- A ranked list of links wore its sourcing on its face. Here are ten places that might know; each carries a name, a reputation, a date; judge them, follow them, notice where they disagree.
- A generated paragraph hides its sourcing by construction. Here is one answer, in one voice, with no seam showing where the knowledge came from, whether it was retrieved from a live source or reconstructed from the model's training, or whether the two were silently blended.
The word for the thing that disappears is provenance — the traceable chain from a claim back to whoever first established it. Provenance is not a nicety. It is the mechanism by which a reader corrects the record. You follow the link, you check the source, you notice that two reputable places tell the story differently, and in that friction you calibrate how much to trust what you just read. A single synthesised answer offers no such friction. It arrives smooth, fluent, and confident, and it does not show you the joins. You cannot follow a link that was dissolved into a sentence.
I want to be exact about the size of this claim, because the failure mode here is to overstate it. Flash is a capable model, and when its answer is grounded in retrieved pages — as search answers often are — it is frequently good, sometimes better than the reading you would have done yourself. The problem is not that the answers are usually wrong. The problem is what happens on the occasions they are. A model has a characteristic way of being wrong: it will state a reconstructed fact with exactly the same fluency as a retrieved one, because to the model there is no felt difference between the two. Memorisation and retrieval produce the same confident prose. And when the wrongness does occur, the reader now has less to catch it with, because the very thing that used to do the catching — the visible source, the competing account — has been designed out of the interface. One model's way of being wrong is, from this month, a great many people's.
A ranked list wore its provenance on its face. A generated paragraph hides its provenance by construction. You cannot follow a link that was dissolved into a sentence.
One translator for everyone
Here is the analogy I would use, once, and then retire before it misleads. Imagine a country in which everyone reads the foreign news through a single translator. The translator need not lie for something to go quietly wrong. Over time, that country's picture of the world acquires the shape of one mind's habits — its emphases and its blind spots, its confidence in the places it should hesitate, the ten percent it consistently renders a little too smoothly. No individual mistranslation is a scandal. The effect is not deception. It is monoculture: a whole population's understanding narrowed to the contours of one intermediary, without anyone choosing that narrowing or being able to see it happen. The risk of a single model answering the world's questions is not that it lies to you. It is that it habituates everyone to the same way of being slightly, invisibly off.
This is where the economics and the epistemics turn out to be one fact seen twice. The reason the answering model is cheap is the same reason a cheap model is the one now shaping a shared understanding of what is true: to answer every question with a language model, the answer has to cost almost nothing, and cheapness selects hard for the model that is confident, brief, and fast rather than the one that shows its work and hedges its uncertainty. Google is building custom silicon — its 'Frozen' chips — precisely to drive that cost down further, so that answering everyone with a model becomes cheaper still. The whole apparatus is being engineered, sensibly and at enormous expense, to make the cheapest possible intermediary the universal one. There is no villain in that sentence. There is just an incentive, running to its conclusion.
The incentive under the answer
It is worth naming the second incentive plainly, because it compounds the first. When Google returned ten links, its business was to send you somewhere else — to a page that carried its own name and, often, its own advertising. When Google writes the answer, its business is to keep you. The 'answer' now competes directly with the sources it is built from; every question resolved on the results page is a click a publisher does not get, a source that does not get the visit that used to sustain it. I state this carefully, because the full consequences are not yet in evidence and I will not pretend they are. But the direction is not in doubt. The arrangement that funded the open web — attention passing through Google to the places that actually did the knowing — is being reorganised so that the attention stops at the answer. A model tuned to say less, hosted by the company that also sells the ads, is now standing between the question and the people who know.
Alphabet reports its quarterly earnings the same week its newest answering model shipped, and the number Wall Street is watching is whether these AI answers erode the search advertising that pays for nearly everything the company does. That is a real question and it will get a real number. But it is not the number I would watch. The one that matters does not appear on an earnings sheet, because we have not yet built the instrument that measures it: what happens to a shared record of what is true when the finding-out is increasingly done for us, by the cheapest model that can keep up, at a scale that forbids anything slower or more careful.
So here is the sharper question, the one underneath the one everyone is asking. The debate has fixed on accuracy — does the AI answer get it right? — and on most queries the honest answer is yes, well enough. But accuracy was never the thing the old interface protected. What it protected was checkability: the reader's standing ability to see where a claim came from and to go argue with it. That is the faculty being quietly removed, not because anyone decided it should be, but because it is expensive to preserve and the model is cheaper without it. The question is not whether Gemini Flash gives good answers. It is whether a species can keep a common footing on the facts when the facts arrive pre-digested, from one mind, priced to move — and when the habit, and the link, that used to take you back to whoever knew it first has been optimised away. We are, as of this month, going to find out.
References
- TechTimes: Google replaced its default search with AI — how to get the blue links back
- 9to5Google: Gemini 3.6 Flash and 3.5 Flash-Lite launch (token and pricing figures)
- Crypto Briefing: Google quietly registers Gemini 3.6 Flash and 3.5 Flash-Lite models
- TradingKey: Alphabet week review — EU hit, Gemini delays, earnings Wednesday
- MarketBeat: Alphabet's AI-spending question looms over Q2 earnings


