AI · Benchmarks

Google delayed its flagship model for missing a coding target. Nobody outside Google can see the target.

The consensus formed in an hour: Google is behind on coding. Behind by how much, measured how, against what? The scores everyone compared it to are mostly run by the companies being scored — and two of them aren't on the leaderboard at all.

The Google logo on the exterior of a building at the company's Mountain View headquarters.

Image: Anthony Quintano / Wikimedia Commons (CC BY 2.0)

Google delayed its flagship model. Bloomberg reported it on Thursday, Alphabet closed down about 4.4 percent, and roughly $200 billion of market value went out the door before dinner. The consensus formed in about an hour, and it was three words long: Google is behind. Maybe. Probably, even — I'll get to why the direction is defensible. But before you repost the take, sit with the question that nobody asked on the way past: behind by how much, measured how, against what? The target that Gemini 3.5 Pro missed is an internal number inside Google. It has never been published. No one outside that building has seen it, and no one is going to. Which means every confident comparison made last week was made against a completely different set of numbers — the public leaderboards — and those have problems of their own that are worth more of your attention than the delay is.

Start with what is actually on the record, because it is narrower than the coverage suggests. Bloomberg reported on 16 July that the flagship is months behind schedule and that Google is taking time to improve the model's capabilities, particularly in coding. The reporting rests on ten current and former employees describing internal frustration, and it attributes some of the drag to structure rather than science: many layers of stakeholders, and separate teams across DeepMind, Cloud, Android and Search building overlapping AI coding tools. It also reports that Google updated Gemini's training data in late June specifically to improve coding, and that the results were disappointing. That last detail is single-sourced. Treat it as reported, not established. Google's on-record statement is deliberately undramatic: the company says it is "currently testing 3.5 Pro, an upgraded Flash model, and other models with partners," and that it is "shipping quickly across a wide range of models while keeping them highly cost-effective for customers." There is no new release date. As of this morning, the model has not shipped.

A note on what I am deliberately leaving out, because it is instructive. Within about a day, the story had been sharpened by downstream outlets into things Bloomberg never reported: a scrapped base model, a ground-up rebuild, a specific abandoned July launch date, named structural failures in recursive tool-calling and SVG generation. I went looking for the source of those claims. They trace to aggregator posts, not to the primary reporting, and then they propagated sideways between sites that cite each other. This is the ordinary metabolism of a tech news cycle: a hedged, sourced report goes in one end, and a set of confident technical specifics comes out the other, having acquired detail it never had. If you saw a version of this story with a launch date in it, you were reading the second thing.

The number Google missed is not a number you can see

Here is the structural problem with the entire week of commentary. "Missed internal targets" is, from the outside, unfalsifiable. Google publishes a great deal about how it gates models on safety — the Frontier Safety Framework, with its critical capability levels across CBRN, cybersecurity, machine-learning R&D and deceptive alignment, and its early-warning evaluations that are supposed to run before external deployment. That framework is real and it is public and it is not what happened here. It governs whether a model is dangerous, not whether it is good. On the capability bar — the actual quality threshold that Gemini 3.5 Pro reportedly failed to clear — Google publishes essentially nothing. No threshold, no eval suite, no score, no methodology.

I want to be precise about the complaint, because it is not an accusation of bad faith. It is a measurement complaint. When a company declines to ship on the strength of a private number, the only honest thing an outside analyst can say is: something inside Google's evaluation returned a result Google didn't like. That is genuinely all the information contained in the event. Everything else that got written last week — the size of the gap, which competitor is ahead, whether this is a blip or a structural decline — was supplied by the writer, not by the evidence. So people reached for the public numbers instead. Reasonable. Let's read them properly, because almost nobody did.

Read the harness before you read the score

Terminal-Bench 2.1 is the benchmark currently doing most of the work in agentic-coding arguments. It measures whether a model, driving a terminal, can actually complete real multi-step engineering tasks. The official leaderboard at tbench.ai does something I wish every leaderboard did: it publishes error bars, and it labels each result according to whether the vendor ran it or an independent party did. Both of those choices turn out to matter enormously, and both get stripped out the moment a number becomes a screenshot.

  • Claude Code with Fable 5 — 83.8% ±1.2 — run by Anthropic
  • Codex with GPT-5.5 — 83.1% ±1.1 — run by OpenAI
  • Terminus 2 with Fable 5 — 80.4% ±1.2 — run independently
  • Claude Code with Opus 4.8 — 78.9% ±1.3 — run by Anthropic
  • Codex with GPT-5.6 Terra — 78.4% ±1.3 — run by OpenAI
  • Terminus 2 with Gemini 3 Pro — 73.9% ±1.3 — run independently
  • Gemini CLI with Gemini 3 Pro — 65.8% ±1.4 — run by Google

Read the first and third rows together. That is the same model — Fable 5, identical weights — scoring 83.8 when Anthropic runs it through Claude Code and 80.4 when an independent harness runs it. Three and a half points of difference produced entirely by the scaffolding around the model rather than the model itself. Now read the last two rows, because they are the ones that should genuinely rearrange your priors. Gemini 3 Pro scores 73.9 on the independent harness and 65.8 through Google's own Gemini CLI. Google's own tooling scores Google's own model eight points below what a neutral harness gets out of it.

Google's own harness scores Google's own model eight points below what a neutral one gets. The vendor-harness effect is not a thumb on the scale. It's noise wearing a brand.

This is the part I would like to nail to the door of every AI newsroom. The reflexive skeptical move — assume vendor-run numbers are inflated — is too simple. Vendor harnesses are not uniformly generous; they are uniformly incomparable. A score on this leaderboard is a joint measurement of a model and the software wrapped around it, and those two things cannot be separated after the fact. Anthropic has invested heavily in its agent scaffolding and it shows up as points. Google's published CLI result makes its model look worse than a third party can. Neither number is a lie. They are just not measurements of the same quantity, and stacking them into a single ranked column produces something that looks like a league table and behaves like a rumour.

And then there are the error bars, which tbench deserves real credit for publishing, and which I have not once seen survive contact with a social post. The top two entries are 83.8 ±1.2 and 83.1 ±1.1. Those intervals overlap comfortably. That is not a first place and a second place. That is a tie, reported as an order, because a ranked list is legible and a confidence interval is not. Rank one and rank two on that board are statistically indistinguishable, and the entire discourse about who leads agentic coding this month is being conducted inside the noise.

The two scores everyone compared Google to aren't on the board

Now the awkward part. The two results most frequently cited last week as evidence of Google's predicament — Moonshot's Kimi K3 and OpenAI's GPT-5.6 — do not appear on the official Terminal-Bench leaderboard at all. Their headline figures are vendor-run and vendor-published.

Kimi K3 arrived in the same news cycle as the Gemini story: 2.8 trillion parameters, the largest open-weight model yet released, with a claimed 88.3 percent on Terminal-Bench 2.1. Read the footnote. That figure is self-reported by Moonshot, produced using the company's own KimiCode harness at maximum reasoning effort. It has not been independently reproduced and it is not on the official board. I am not calling it wrong — I have no basis to. I am saying it is currently a claim, and it is being compared against numbers generated under different conditions. To Moonshot's genuine credit, this is a claim with an expiry date on it: the weights are due on 27 July, at which point anyone with sufficient hardware can check the homework. That is the difference between an open-weights claim and a closed one, and it is the strongest argument for open weights that exists.

GPT-5.6, which went generally available on 9 July, is messier still. There are three defensible numbers in circulation, all of them honestly labelled somewhere and all of them describing "GPT-5.6 on Terminal-Bench 2.1." OpenAI self-reports 91.9 percent for Sol Ultra — but Sol Ultra is a high-compute parallel-agent mode, not a like-for-like single-agent run, which makes putting it in a column next to anything else a category error. Artificial Analysis, running independently, gets 89.5 percent for Sol at extra-high reasoning effort. The official board lists Terra via Codex at 78.4. That is a spread of more than thirteen points, and which value you drop into your comparison quietly determines your conclusion before you have made an argument. Most of last week's takes picked the highest one.

So let me credit the one measurement in this story that is built properly. Arena.ai's Frontend Code Arena runs blind pairwise human voting — real people choosing between two anonymised outputs, scored by a third party with no model in the race. Kimi K3 took the top spot there at 1,679 Elo, ahead of Fable 5 at 1,631 and GPT-5.6 Sol at 1,618, jumping seventeen places from its predecessor's rank, with a 76 percent pairwise win rate. That is a real, independently administered result and I am happy to report it as one. The footnote, since I am obliged to read it out: it rests on 1,757 valid votes. That is a respectable sample and not an enormous one, and an Elo gap of roughly forty-eight points on that many comparisons carries a confidence interval that nobody screenshotted either.

What Gemini actually scores

With the harnesses accounted for, where does Google genuinely sit? Gemini 3 Pro posts 76.2 percent on SWE-bench Verified and 54.2 percent on Terminal-Bench 2.0. That second figure needs a warning label welded to it: 2.0 is not 2.1, the task sets differ, and I have watched people put those two numbers in the same column this week as though the version suffix were decoration. Gemini 3.5 Flash — the model that did ship, back at I/O in May — self-reports 76.2 percent on Terminal-Bench 2.1, and Google claims it beats the older 3.1 Pro on coding and agentic work at around 40 percent lower cost and four times the speed.

Put the defensible figures side by side and the honest summary is this: Google's best publicly benchmarked coding result sits somewhere around eight to twelve points below the frontier that its competitors claim, and the model that was supposed to close that gap is precisely the one that slipped. The direction of the consensus is right. The magnitude is far softer than the confident posts implied, a meaningful fraction of the apparent gap is harness rather than model, and the top of the comparison set is made of vendor claims that no independent party has reproduced. "Google is behind" survives scrutiny as a direction. It does not survive as a measurement.

There is also a piece of on-the-record context that deserved more attention than the stock move got. At I/O in May, Sundar Pichai said plainly that Google was "a bit behind" on agentic coding, and attributed it to a lack of developer-facing products generating competitive training data. That is the single most useful sentence anyone at Google has said about this, it was said publicly two months before the delay was reported, and it describes a data problem rather than a modelling one. If the late-June training-data update was the attempt to fix exactly that, and it disappointed, then the delay is not a surprise. It is a confirmation.

Why the long tasks are where it breaks

One more thing worth understanding, because it explains why a model can look excellent on a benchmark and still fail the work. Per-step accuracy compounds multiplicatively, and the arithmetic is unforgiving in a way that intuition does not prepare you for. A model that is 98 percent reliable on each individual step of a twenty-step dependent task has roughly a one-in-three chance of botching at least one of them. Drop per-step accuracy to 80 percent across ten dependent steps and the probability of a clean run through is about 10.7 percent. Roughly one attempt in nine. The same system that writes a flawless function on demand fails a two-hour engineering task nine times out of ten, and no single-step benchmark you can name predicts that failure.

At 80 percent accuracy per step, a ten-step task completes cleanly about one time in nine. That is the distance between a model that codes and a model that finishes.

This is an active research area rather than folk wisdom. Work published this year analysing more than 3,100 agent trajectories across model families finds that long-horizon breakdown is not simply a lower success rate — it is a structural shift in the composition of failures, with different things going wrong at length than go wrong in short tasks. Related work identifies error propagation as the primary reliability bottleneck, with early mistakes cascading through everything downstream, and finds that models remain linguistically fluent while diverging functionally, which is precisely the failure mode that is hardest to catch automatically. The model keeps sounding correct long after it has stopped being correct. If that is the regime Gemini 3.5 Pro is stuck in, then "improve coding" is not a tuning pass before launch. It is the hard open problem in the field, and Google would not be the only lab that has not solved it.

What the delay actually supports

So: verdict. Two readings were on offer last week and both are wrong. "Google is losing the AI coding race" is overstated — it converts an unpublished internal number and a set of incommensurable public ones into a confidence nobody has earned. "Nothing happened" is also wrong; months of slippage on a flagship, and $200 billion of market value, are not nothing. What the evidence actually supports is narrower and considerably more interesting than either. Google ran an evaluation, did not like what came back, and declined to ship. Every other lab in this story spent the same fortnight shipping a number.

I would rather have Google's unpublished failing score than most of the published passing ones, and not because I think Google is more honest than anyone else. It is because a company that ships on a target it missed has told you nothing at all, while a company that eats a four percent drawdown rather than release something it privately grades as inadequate has told you that at least one real threshold exists somewhere in its process. That is a low bar. It is still a bar, and it is more than a self-run harness at maximum reasoning effort can clear.

The date to watch is not Google's release, which has no date. It is 27 July, when Kimi K3's weights land and that 88.3 becomes something anyone can check rather than something everyone has to repeat. I intend to check it. Until then, the agentic coding leaderboard is a set of vendor claims with a ranking painted over the top, a couple of genuinely independent results doing all the load-bearing work, and error bars wide enough to swallow the first two places. The chart went up all week. Look harder.

References

  1. Bloomberg: Google Gemini launch delayed as tech falls short of internal goals
  2. 9to5Google: Gemini 3.5 Pro delays due to coding performance
  3. Google: introducing Gemini 3.5 (I/O 2026)
  4. Terminal-Bench 2.1 official leaderboard (tbench.ai)
  5. Artificial Analysis: independent Terminal-Bench v2.1 evaluations
  6. Google DeepMind: strengthening the Frontier Safety Framework
  7. The Long-Horizon Task Mirage? Diagnosing where and why agentic systems break (arXiv)
The Friday Brief

One email. Every Friday.

The week's machines, money, and people — in under five minutes.