OpenAI picked Nvidia's earnings day to attack the one number holding up the trade
OpenAI's Jalapeño chip claims to beat Blackwell on work per watt. The benchmark is arguable. The target — Nvidia's roughly 75% margin, and how much of the index quietly rests on it — is the real position.

Image: 极客湾Geekerwan / Wikimedia Commons (CC BY 3.0)
Two things happened on Wednesday, and only one of them was on the calendar. Nvidia reported its second quarter after the close — the most-watched print in the market, roughly $92 billion of revenue that the entire AI trade had spent a month pricing in. And a few hours earlier, its largest customer released the first benchmarks for a chip designed to need less of it. OpenAI did not schedule Jalapeño's debut for a random Wednesday. It scheduled it for this one.
The number everyone repeated by lunchtime was the headline OpenAI wanted repeated: its custom inference processor, built with Broadcom, delivered 1.5 to 1.9 times more AI work per watt than Nvidia's flagship Blackwell systems, and up to several times lower latency, in the first published tests. A 700-watt chip beating a 1,400-watt one. Faster, on the chatty, low-latency traffic that a consumer chatbot actually generates. It is a good headline. It is also, for the reader trying to understand the position rather than the press release, the least important sentence in the story.
Here is the number that matters, and it is not on Jalapeño's spec sheet. It is on Nvidia's income statement: a gross margin of roughly 75 percent. Three quarters of every dollar Nvidia books in sales is gross profit. That figure — not the revenue line, not the market cap — is the load-bearing wall of the entire AI equity trade. OpenAI just spent an announcement, on Nvidia's earnings day, telling the market that the wall has a crack in it. Whether it does is a longer question than the benchmark can answer. But you should understand what is being aimed at before you decide whether the shot landed.
The Nvidia tax, stated as a margin
Start with the mechanism, because the mechanism is where the money is. A company that sells a commodity earns a thin margin; a company that sells the one thing everyone needs and no one else makes earns a fat one. Nvidia's roughly 75 percent gross margin is not the reward for building a good chip. It is the price of being the only practical way to run frontier AI at scale — the toll every lab, cloud and startup has paid because the alternative was not to compute at all. Inside the industry there is a phrase for the gap between what a Nvidia system costs to make and what it costs to buy. They call it the Nvidia tax. On the income statement, the tax has a name too. It is called gross margin.
Now separate the workload, because OpenAI did. There are two jobs an AI chip does. Training builds the model — a brutal, once-in-a-while, capital-intensive job where Nvidia's lead is real and, for now, close to unassailable. Inference runs the model — the billion small, repetitive, latency-sensitive requests that happen every time someone actually uses the thing. Training is the cathedral. Inference is the turnstile, and the turnstile is where the recurring money is. Jalapeño does not touch training. It is built for one job: to run models that already exist, cheaply, at the turnstile. Broadcom's chief executive put a number on the ambition to Reuters and Bloomberg — roughly half the inference cost per token. Half the tax, on the half of the workload that repeats forever.
That is the position OpenAI is really taking. Not 'we built a faster chip.' The claim underneath the benchmark is: the most profitable, most repeatable slice of Nvidia's business is the slice that a determined customer can carve off and do for itself. If that is true, Nvidia's margin is not a moat. It is a temporary toll on a road someone is quietly building a bypass around.
Nvidia's 75 percent gross margin is the load-bearing wall of the AI trade. OpenAI just told the market, on earnings day, that the wall has a crack in it. — The position under the benchmark
What OpenAI actually owns
Before pricing the crack, price the claim honestly — the way you would if you had to put money behind it rather than a headline. And the honest read is that the benchmark is softer than the coverage suggests, in ways that matter to the trade.
The tests ran on SemiAnalysis's public InferenceX suite, across three open models. But OpenAI supplied the numbers. SemiAnalysis says it observed selected runs and did not itself execute the full suite — which is a long way from an independent, reproducible benchmark with raw artifacts and matched system telemetry. Jalapeño was measured against the GB200 and GB300, not against Nvidia's newer Vera Rubin parts, which have only just begun to ship. And it is an inference-only chip, so the comparison deliberately excludes the training workloads where Nvidia's lead is not in question. A vendor released favorable numbers, on its rival's biggest day, under conditions it controlled. That does not make the numbers false. It makes them a marketing document with a benchmark attached, and a careful reader discounts it accordingly.
Then look at what OpenAI owns, versus what it rents. It owns a design — a drawing of a chip. It does not own a fab; the silicon is made by TSMC. It does not own the manufacturing partnership; that is Broadcom, whose other customers are queued behind the same insatiable demand. It does not own the deployment; Jalapeño is slated to start rolling out inside OpenAI's own infrastructure and at partners including Microsoft by the end of 2026, with the real volume ramp not arriving until 2027. And it does not own the capital or the power; the whole thing sits inside a buildout measured in gigawatts and financed by other people's balance sheets.
- What the benchmark establishes: a custom inference ASIC can, on models OpenAI chose, post better work-per-watt than a general-purpose GPU. This was already true in principle — Google's TPU, Amazon's Trainium and Meta's MTIA proved the pattern. Jalapeño is a data point, not a revelation.
- What it does not establish: that the advantage holds against Nvidia's current-generation parts, that it survives independent testing, or that OpenAI can manufacture at the volume its own compute roadmap requires.
- What it changes today: nothing on the income statement. Jalapeño ships in volume in 2027. Nvidia's margin is a 2026 fact.
- What it changes tomorrow: the negotiation. A customer that can credibly build its own inference chip is a customer that pays less for yours. The threat to build is worth money even if the chip never dominates.
That last point is the one to hold onto, because it is the one that is real now. OpenAI does not need Jalapeño to beat Nvidia. It needs Nvidia to believe Jalapeño might. The benchmark is a letter to a supplier as much as a spec sheet — and it was mailed, first-class, on the day Nvidia had the market's full attention.
The concentration no one is pricing
Now the part that has nothing to do with chips and everything to do with your retirement account. Nvidia is the largest single weight in the S&P 500 — around 7 to 8 percent of the index, depending on the day. That means every passive fund tracking that index, every default 401(k) allocation, every 'safe' broad-market position, is a concentrated bet on one company. Most of the people holding that bet do not know they are holding it. They think they own the market. They own, to a degree without modern precedent, one stock.
And that one stock is valued the way it is because the market has priced its margin as durable. A 75 percent gross margin capitalized into a multi-trillion-dollar valuation is not a bet that Nvidia sells more chips. It is a bet that Nvidia keeps 75 cents of every dollar for years — that no customer, no rival, no shift in the workload ever competes the tax down. Concentration and margin assumption are the same position viewed from two angles. The index leans on Nvidia; Nvidia leans on its margin; the margin leans on the absence of exactly the kind of thing OpenAI announced on Wednesday.
I want to be precise about what I am and am not saying. I am not saying the margin collapses. Nvidia has answered every skeptic for three years by simply growing into the doubt, and Blackwell demand, by every account including its own earnings, remains extraordinary. The training moat is intact. The software lock-in is real. A custom chip that ships in volume in 2027 does not dent a margin booked in 2026. The bull case is not stupid; it has been right repeatedly, and being early to call it is its own kind of wrong.
The downside case
But risk is not what you expect to happen. Risk is what your position cannot survive. And the position embedded in the index — long Nvidia's margin, at scale, by default — has a specific vulnerability that Wednesday illuminated. It is not that Nvidia sells fewer chips. It is that the most profitable part of what it sells, inference, is the part its own largest customers have both the incentive and, increasingly, the capacity to take in-house. Google did it. Amazon did it. Now the customer that defined the demand curve, the one whose name is synonymous with the boom, has published a benchmark saying it can too.
You do not need the margin to fall to 40 percent for the trade to hurt. You need the market to stop believing it will stay at 75. Multiples are not paid on this year's profit; they are paid on the durability of it. The day the consensus shifts from 'Nvidia keeps the tax forever' to 'Nvidia keeps the tax on training and competes for inference,' the number that reprices is not revenue. It is the multiple on the margin — and because the index is concentrated in that one name, the repricing does not stay contained. The same mechanism that carried the index up on the way in, everyone owning the same thing and calling it prudence, has no brakes on the way out, because the buyers and the believers are the same people.
That is why the timing was the message. OpenAI could have released these benchmarks any day. It chose the one day the whole market was already looking at Nvidia's margin, and it used the attention to introduce a doubt. The chip may or may not matter in 2027. The doubt is priced starting now — a little at first, in the questions on the next few earnings calls, in the way analysts start to model inference and training as separate businesses with separate margins. Whether it compounds is the only question in the AI trade that is actually worth watching. Jalapeño is a chip. What OpenAI aimed it at is a number, and that number is holding up more of the market than most of the people exposed to it realize.
References
- OpenAI — OpenAI and Broadcom unveil LLM-optimized inference chip
- Tom's Hardware — OpenAI's 700W Jalapeño ASIC outpaces 1,400W Nvidia flagship in first published benchmarks
- CNBC — OpenAI's Jalapeño chip brings new 'threat' to Nvidia margins
- SemiAnalysis — OpenAI Jalapeño: Better Than Nvidia Blackwell
- The Motley Fool — Nvidia earnings on August 26: what history tells us
- Broadcom — OpenAI and Broadcom unveil LLM-optimized Intelligence Processor (press release)


