Benchmarks

OpenAI says its model just got 14 times faster. Every number in that sentence was picked by the company selling it.

"Up to 14x," "750 tokens a second," a preview you can't get into: the speed race has produced a set of figures with no published method and no way to check them — and speed is quietly being read as if it were quality.

Cerebras wafer-scale processor, the hardware OpenAI says powers its new Ultrafast inference tier.

Image: Cerebras Systems

OpenAI previewed a new service tier this week called Ultrafast, and the number that travelled was the big one: GPT-5.6 Sol, the company's flagship model, running up to 14 times faster than standard, peaking at around 750 output tokens a second, powered by Cerebras's wafer-scale hardware. Fourteen times. The posts wrote themselves. Here is the question almost nobody asked before reposting the figure: 14 times faster than what, measured on which task, at what prompt length, at what batch size, and giving the same answers as the slow version or different ones. And then the more awkward question underneath all of those: can anyone outside OpenAI actually check the claim right now. The answer to the last one is no, and that is where I want to start, because a number you cannot reproduce is not a measurement. It is a demo with a decimal point.

I want to be careful here, because the skepticism is about the number, not the achievement. Wafer-scale inference is a real thing and Cerebras genuinely does move tokens at speeds conventional GPU clusters struggle to match; 750 tokens a second is plausible, not physically outlandish. This is not a fraud story. It is a measurement story, which is quieter and more common: a set of impressive figures released with no method attached, into a moment when the industry has started selling speed the way it used to sell quality, and readers have started reading the two as if they were the same thing. They are not, and the gap between them is the whole of this piece.

"Up to," and the baseline nobody names

Start with the two smallest words in the claim, because they are doing the most work: up to. "Up to 14x" is not a measurement of typical performance; it is a ceiling, the single best ratio observed under the single most favourable condition, reported as if it were the number you will see. Every honest engineer knows the real distribution sits somewhere below it, often well below, and "up to" is precisely the phrase you reach for when you would rather not publish the average. Ask what the typical speed-up is on a normal request and there is no figure, because none was released. There is a peak and a multiple, and the multiple attaches to the peak.

Now the baseline. Fourteen times faster than what? The comparison is to OpenAI's own Standard tier of the same model. Read that again, because it is the tell: the company is measuring its fast product against its slow product, both of which it prices and controls, and reporting the ratio between them. That ratio can be perfectly true and still tell you almost nothing, because OpenAI sets both ends of it. It says nothing about whether 750 tokens a second is fast relative to what the same Cerebras silicon does for an open-weights model you could benchmark yourself, or relative to the other speed-specialist services, or relative to what a well-configured GPU deployment hits on a good day. A self-referential multiple — our fast thing versus our slow thing — is the easiest number in the world to make large, because you can always make the denominator slower. Compared to what is the first question you ask any benchmark, and here the honest answer is: compared to another number the same company chose.

Three different speeds, one headline

There is a deeper confusion buried in the word "faster," and it survives because most coverage never separates the three things it can mean. When people say a model is fast they are usually blurring together three distinct measurements that do not move together and often trade against each other.

  • Time to first token — how long you wait after hitting send before anything appears. This is latency, and it is what makes a chat feel responsive. A wafer-scale engine can be brilliant here or ordinary, depending on how the request is scheduled.
  • Throughput — tokens per second once the stream is flowing. This is the 750 figure. It matters most for long generations and for agents that produce reams of text, and it is the number Ultrafast is built to win.
  • End-to-end task time — how long the whole job actually takes, including the model thinking, calling tools, and any reasoning steps before it writes a word. For a modern reasoning model this can dwarf the raw token stream, which means a 14x throughput gain can shrink to a far smaller wall-clock gain on the task you actually care about.

"Fourteen times faster" almost certainly refers to the second of those, throughput, because that is the one wafer-scale hardware wins most decisively. But it will be read as the third — my work gets done fourteen times sooner — and those are not the same claim. If a Sol request spends most of its time reasoning before it emits tokens, speeding up the token emission fourteenfold moves the part of the clock that was already small. I cannot tell you by how much, and that is exactly the point: neither can you, from what was published, because the one measurement that would answer it — end-to-end task time on a defined workload — is not in the announcement.

A self-referential multiple — our fast thing versus our slow thing — is the easiest number in the world to make large. You can always make the denominator slower. — On what "14x faster" actually compares

The error bar that matters: same speed, same answer?

Here is the number I most want and cannot find, and it is the one that should worry anyone about to route real work through Ultrafast. Does the fast Sol give the same answers as the standard Sol? Speeding a model up on exotic hardware is not always free of consequences. It can involve running at lower numerical precision, changing how requests are batched, or serving a variant tuned for the accelerator — and any of those can, in principle, nudge the outputs. Sometimes the nudge is invisible. Sometimes it shows up as slightly worse instruction-following, or a few more points of hallucination, or degraded performance on the long, hard, multi-step problems where flagship models earn their keep. OpenAI did not publish a quality comparison between the tiers. There is no table showing that Ultrafast Sol scores the same as Standard Sol on any evaluation at all.

Maybe they are identical. Maybe the wafer-scale path is bit-for-bit the same model and the only thing that changed is the clock. I genuinely do not know, and I am saying so loudly, because the absence of the comparison is itself information. When a company has a clean result it publishes the table; the table is cheap and it flatters you. When the comparison is missing, the safe assumption is not that it would have looked great. At minimum, "same model, just faster" is a claim of no quality regression, and a claim of no regression with no benchmark behind it is a hope wearing the clothes of a fact. If you are moving production traffic to the fast tier, the question to put to your account manager is not how many tokens a second. It is: show me Ultrafast Sol and Standard Sol on the same held-out set, with error bars, and let me see they land in the same place.

A preview you cannot run is a claim, not a result

All of which runs into the wall that makes every figure above unverifiable today: Ultrafast is a waitlist-gated preview. No general-availability date. No published price. Available at first only through the API, only for Sol, only to selected customers, expanding "as capacity grows." I understand why — wafer-scale capacity is scarce and you ration a scarce thing — and to OpenAI's credit the announcement says all of this plainly rather than pretending the tier is open. That candour is real and I will credit it. But it also means the 14x and the 750 are, for now, numbers the vendor reports about a product almost no one can independently touch. In benchmarking that is the least verifiable category there is: a result you are asked to believe because the source is confident, not because you can re-run it. The proper name for a number you cannot reproduce is not a lie and it is not a measurement. It is a claim, and claims are graded on the track record of whoever makes them, not on how round they are.

This is not unique to OpenAI, which is what makes it worth flagging now rather than later. The whole field has spent this year turning inference speed into a marketed axis. There is a Standard tier, a Fast tier, and now an Ultrafast tier; speed has become a product with a pricing page, the way storage or bandwidth is, and every one of the competitive multiples flying around — this service is Nx faster than that one — traces back, when you follow it, to a figure the seller produced under conditions the seller chose. That is not automatically dishonest. It is just unaudited, and unaudited numbers have a way of drifting upward, because the incentive all points one direction and nobody is paid to publish the confidence interval.

What the number actually supports

So what can you take from this week's announcement, stripped of the multiple? A fair amount, actually, as long as you keep it inside its evidence. Ultrafast is real and represents a genuine engineering push: OpenAI is putting its flagship model on Cerebras's wafer-scale hardware and can, under favourable conditions, stream tokens dramatically faster than its standard GPU path. For workloads dominated by long generations — bulk drafting, code emission, high-volume agents that produce a lot of text — that is a meaningful improvement, and speed at that scale is not a vanity metric; latency and throughput are real costs and real user experience. If you live in those workloads, this is worth your attention.

What the number does not support is the reading it will mostly get: that GPT-5.6 Sol is now, in some general sense, fourteen times better, or that your particular job will finish fourteen times sooner, or that the fast version is the same model in every way that matters. "Up to 14x" is a throughput ceiling, measured against OpenAI's own baseline, on undisclosed conditions, for a model you cannot yet test, with no quality comparison attached. That is not nothing. It is also not the sentence people are repeating. The two figures worth waiting for are the ones that were left out: a same-workload, same-quality comparison between the tiers, and a price. Those are the numbers that turn speed from a slide into a fact you can plan around — and until they arrive, the honest thing to write above the 14x is the phrase the industry keeps hoping you will forget to add. Reported by the vendor. Not independently verified. Look harder.

References

  1. The Decoder — GPT-5.6 Sol goes 14x faster as OpenAI launches Ultrafast mode powered by Cerebras
  2. OpenAI — Advancing the price-performance frontier with GPT-5.6
  3. Digital Applied — GPT-5.6 Sol Ultrafast: OpenAI previews 14x inference
  4. OpenRouter — GPT-5.6 Sol: API pricing and benchmarks
  5. EdenAI — GPT-5.6 Sol: benchmarks, pricing and API access guide
The Friday Brief

One email. Every Friday.

The week's machines, money, and people — in under five minutes.