Custom silicon

Meta's new AI chip isn't built to beat Nvidia. It's built to survive the power bill.

Meta starts manufacturing its Iris accelerator this month. The only number that decides whether it matters is the cost of running 14 gigawatts of compute — roughly half the electricity of New York City.

Rows of server racks in a data center hall

Image: "Datacenter Server Racks" by Carl Lender / Wikimedia Commons, CC BY 2.0

The most important number in Meta's new AI chip is not printed on the chip. It is printed on the electricity bill.

Meta will begin manufacturing its third in-house accelerator this month. The chip is called MTIA v3, codenamed Iris, designed with Broadcom and fabricated on TSMC's 3-nanometer process, built to handle both training and the low-latency inference that serves a recommendation feed. That is the spec sheet, and the spec sheet is the least interesting part of the story. Iris is not a faster chip than Nvidia's. By most accounts it does not need to be. It needs to be cheaper to run, and Meta has put a figure on how much cheaper: an estimated 40 to 44 percent reduction in total cost of ownership when the work runs on its own silicon instead of someone else's.

Hold that percentage next to the thing it is being applied to. Meta has said it wants to reach roughly 14 gigawatts of computing capacity by 2027. Fourteen gigawatts is not a data-center number in the way people usually mean it. It is a grid number. It is the kind of figure you find in the annual report of a national utility, not the roadmap of a social network. And that is the real reason a company that could simply keep buying the best chips on the market is instead building a worse one on purpose.

What 14 gigawatts actually costs

Start with the physical plant, because the physical plant is where the money goes. Meta's Prometheus cluster in New Albany, Ohio, is coming online this year as a roughly 1-gigawatt supercluster. Its Hyperion campus in Louisiana is the larger bet: a build that Mark Zuckerberg has described as scaling to around 5 gigawatts of compute in the years after 2027, with an interim target of 1.5 gigawatts by the end of that year. To feed Hyperion, Meta has arranged for something on the order of 10 gas-fired power plants with about 7.5 gigawatts of generating capacity — an addition that, by itself, would represent more than a 30 percent increase to the entire grid capacity of Louisiana, before counting the renewable and storage capacity Meta has also agreed to help fund.

Put plainly: at peak, one of Meta's campuses will draw about half as much electricity as the city of New York. That is the denominator Iris is aimed at. When a company is provisioning power on the scale of a mid-sized country, the difference between a chip that is efficient for your specific workload and one that is merely powerful stops being a benchmark footnote and becomes the difference between two utility bills, compounding every hour of every day for the life of the hardware.

This is the arithmetic that never makes the keynote. A frontier GPU is sold on peak performance — the number that wins the slide. But the number that wins the fleet is cost per unit of useful work, sustained, at the temperature and utilization of a real building. Total cost of ownership folds in the price of the silicon, the margin the vendor takes on top of it, the power to run it, the cooling to keep it alive, and the years it stays in service. Cut any one of those and the headline performance barely moves. Cut all of them for the exact jobs you run most — ranking, ads, Llama inference — and you can field a chip that loses on paper and wins on the invoice.

A frontier GPU is sold on the number that wins the slide. Meta is building for the number that wins the fleet: cost per unit of useful work, sustained, at the temperature of a real building.

The Nvidia margin is a line item you can delete

There is a reason every hyperscaler with the engineering budget to do it has ended up here. When you buy an accelerator from Nvidia, you pay for the transistors and you pay for Nvidia's gross margin, which has been among the widest in the industry. For a buyer purchasing a handful of servers, that margin is the price of not having to design your own chip. For a buyer provisioning 14 gigawatts, it is a recurring tax large enough to justify building a chip company inside your company. Meta is, on this measure, in catch-up mode with its peers — Google put custom tensor chips into its data centers back in 2014 — and the estimated 40-plus percent TCO improvement is, in large part, the sound of that margin being removed from the bill.

The honest way to read the 40-to-44 percent figure is as an estimate produced by the company that benefits from it, modeled against its own workloads, before the first wafer has run at volume. It is not a measured result. First silicon rarely hits its modeled economics on the first pass; yields disappoint, a workload shifts, a cooling assumption turns out optimistic. The number is a target, and the interesting question is not whether Meta hits it exactly but whether it lands close enough, often enough, to justify the pace it has set for itself.

That pace is the part worth watching. Meta's MTIA program is now reportedly aiming to ship a new generation roughly every six months through the end of the decade — Iris this year, then successors reportedly codenamed Santa Barbara, Olympus, and Universal Core across 2027 and 2028. A six-month cadence compresses what is normally a 12-to-24-month design cycle into half the time. It is the schedule of a company that has decided the compute build-out will not wait for the chip, and that the chip must therefore learn to run at the speed of the build-out. That is an aggressive bet, and aggressive bets on silicon schedules are exactly the kind that slip.

The cost you cut inside is a cost someone else pays

Here is where the story usually stops and where it should not. Total cost of ownership is an internal number. It measures what the compute costs Meta. It does not measure what the compute costs the grid it plugs into, and those are not the same ledger. When a single campus adds gas-fired generation equal to a third of a state's grid, the cost of that power — its price, its emissions, the water drawn to cool the servers, the strain on transmission — does not vanish because a custom chip made the workload cheaper for the company running it. It moves. It lands on ratepayers, on utilities negotiating who pays for new lines, on communities that had no vote on becoming the substrate for someone else's inference.

This is the pattern anyone who has covered energy outside Silicon Valley learns to look for. Efficiency at the device rarely means less total energy; it usually means more, because cheaper compute gets used more. A chip that cuts the cost of a token by 40 percent does not lead a company to run 40 percent fewer tokens. It leads it to run more of them, which is the entire point of building the chip. Iris is not a conservation story. It is a scaling story wearing the language of efficiency, and the two look identical right up until you read the utility's interconnection queue.

None of that makes the chip a bad decision. From Meta's chair it is close to an obvious one: if you are going to spend on the scale of a national utility, you should own the most expensive, most repeated component in the stack rather than rent it. But the deployment that follows is not a private matter settled inside a data center. It is a claim on real generation, real transmission, and real water, in real places — Ohio, Louisiana — that will feel the draw long before they see the benefit.

What decides whether Iris matters

Strip away the codenames and the process node and the roadmap, and a custom accelerator succeeds or fails on a short list of unglamorous questions:

  • Cost per token at scale. Not peak performance — the sustained, all-in cost of a unit of useful work on Meta's own workloads, measured after the chip has run at real utilization for a year. That is the number the 40-to-44 percent estimate is promising, and the only one that validates it.
  • Whether the estimate survives contact. TCO models are built on assumed yields, assumed utilization, and assumed power prices. All three drift. The chip does not have to hit the model exactly; it has to miss it by little enough that the economics still beat buying Nvidia.
  • Whether the grid can actually deliver 14 gigawatts. A chip that is cheap to run is worthless if the power to run it arrives late or over budget. The binding constraint on Meta's compute may not be the silicon at all — it may be the interconnection queue and the gas-turbine lead times.
  • Whether the six-month cadence holds. The TCO advantage assumes a steady march of better silicon. If Santa Barbara or Olympus slips, the fleet ages, and an aging custom chip loses to the merchant part it was meant to replace.
  • Who ends up paying for the power. The internal cost curve and the external cost curve are different lines. The first is Meta's to optimize. The second is a bill sent to a grid, and it comes due whether or not the chip hits its number.

The temptation with a story like this is to grade it on the chip — the node, the transistor count, whether it beats a B300 on some benchmark it was never designed to win. That is the wrong scorecard. Iris was not built to win a benchmark. It was built to change the denominator on a power bill that is about to be one of the largest in American industry. Watch the cost per token across 14 gigawatts, and watch the invoice the grid sends back. Those two numbers, not the spec sheet, will tell you whether the most expensive chip Meta ever made was the one it decided to stop buying.

References

  1. Electronics Weekly — Meta prepares to fab its third custom datacentre chip
  2. Data Center Dynamics — Meta could start production of Iris AI chip in September
  3. Data Center Frontier — Ownership and power challenges in Meta's Hyperion and Prometheus data centers
  4. Fortune — Meta orders 10 gas-fired power plants for its Hyperion AI campus in Louisiana
  5. IEEE ComSoc Technology Blog — Meta's Iris AI chip for MTIA
The Friday Brief

One email. Every Friday.

The week's machines, money, and people — in under five minutes.