Silicon

Etched drew the transformer into the silicon itself. Investors just paid $21 billion to bet the math never changes.

Its chip can run one kind of AI model and nothing else. That is the whole product — and, if the architecture that rules the field ever shifts, the whole risk.

An Etched inference rack, the company's transformer-specific AI hardware

Image: Etched

On most of the chips running artificial intelligence today, the mathematics is a set of instructions. A graphics processor is a room full of general-purpose arithmetic units, and the specific pattern of a transformer — the attention step, the projections, the feed-forward layers that turn a prompt into the next word — is orchestrated across them in software, assembled and torn down billions of times a second. On the chip a startup called Etched has begun shipping, that pattern is not orchestrated. It is drawn. The operations a transformer performs are laid into the physical floorplan of the silicon, fixed in the arrangement of the transistors themselves, so that the math a GPU runs as a program is, here, the shape of the metal. A chip like that can do one thing. The wager Etched has just raised $700 million against is that the one thing is the only thing that will matter.

The raise, announced this week, values the company at $21 billion. The number is worth pausing on, because of how fast it arrived. In December Etched was valued at around $5 billion. In late July it closed a $300 million round at $10.3 billion. Weeks later, a new $700 million round led by the quantitative trading firm Jane Street — with Kleiner Perkins, Sequoia, Andreessen Horowitz, Tiger Global, Bain Capital Ventures and Blackstone alongside — has roughly doubled that figure again. A company founded in 2022 by two Harvard dropouts, which emerged from stealth barely two months ago, is now valued above most of the established chip firms whose products it hopes to displace. Valuations move on stories. This one moved on a single, specific story about the shape of a chip.

What Sohu is, and what it threw away

The chip is called Sohu, and understanding what it is requires understanding what its designers deleted. A modern GPU is a generalist by construction. It spends a great deal of its silicon on flexibility — on the ability to run any computation a programmer might throw at it — because its makers cannot know in advance what the workload will be. Etched made the opposite bet. Its engineers assumed the workload. They took the transformer, the architecture underneath essentially every large language model since 2017, as a fixed fact, and then stripped out everything on the chip that exists to run anything else.

The die area freed by that deletion was spent on the operations a transformer actually performs. Attention, the projections, the feed-forward math are hard-wired into the layout; the arithmetic a GPU choreographs in software is, on Sohu, the floorplan. The chip is fabricated on TSMC's four-nanometer process — the same leading-edge Taiwanese line that builds the advanced parts for Apple and Nvidia — and each one carries 144 gigabytes of high-bandwidth memory. The company has described two custom pieces: a prefill design run at low voltage to pack in more transistors without cooking itself, and a decode design that lets a cluster of chips draw on one fast, shared pool of memory. None of this is exotic in its ingredients. What is unusual is the conviction: to build a chip this specialized, you have to believe the workload will not move.

The claim, and the caution it deserves

The performance Etched claims for that specialization is enormous, and it should be read with the caution any first-party benchmark deserves. A server of eight Sohu chips, the company says, can generate more than 500,000 tokens a second running Llama-70B, a common open model; an equivalent server of eight Nvidia H100 GPUs, it says, manages roughly 23,000. That is not an incremental gain. It is a claimed order of magnitude and then some.

It is also, at this writing, a number produced by the company selling the chip. Sohu is early silicon — first-pass, in the industry's phrase, meaning the very first version to come back from the fab and actually work, which is itself a real achievement and not the same thing as a mature product. There are no independent benchmarks. There is, as of this week, exactly one customer with the hardware in hand. Every figure above should be held at the arm's length that distance implies.

A GPU is a hedge. An ASIC is a conviction. Etched has cast its conviction in four-nanometer silicon — and the bet is that the transformer is a destination, not a phase.

Why a trading firm went first

That one customer is the detail that tells you the most, and it is not a coincidence that it also led the round. Jane Street is a quantitative trading firm, a business whose profits are made or lost in the microseconds it takes to turn information into a decision. To a firm like that, the speed of inference is not a feature to admire; it is money, measured directly. Etched says it has shipped its first rack to Jane Street, which has installed the system in its own data center and begun running its own workloads across it — 'pleased with the early results,' in the firm's careful phrase.

A quant shop is close to the ideal first buyer for a chip like this. It has a real, punishing workload, the in-house engineering to validate the hardware honestly rather than take the datasheet on faith, and a balance sheet that makes a bet on unproven silicon survivable if it fails. It is also, as the lead investor in the round, not a disinterested referee. The early results are encouraging, and they are being reported from inside the tent. Both things are true at once, and a careful reader holds them together.

The pattern that says yes, and the graveyard that says wait

There is a well-worn pattern under all of this, and it cuts both ways. Specialized silicon beating general-purpose silicon at a fixed task is one of the oldest moves in the industry. It is what happened to Bitcoin mining, which fled the GPU for purpose-built chips that did nothing but hash, and did it so much more efficiently that general-purpose parts left the field entirely. It is the logic behind Google's own TPU, built because a company running one kind of computation at vast scale can justify the enormous fixed cost of a chip that does only that. When the workload is stable and the volume is high enough, the generalist loses to the specialist. That is the strongest version of the Etched thesis, and it is a serious one.

But the same history has a second half, and it reads like a graveyard. The field is littered with AI-chip startups — well-funded, technically credible, celebrated in their moment — that bet on a particular shape of the workload and were overtaken when the shape moved. A chip takes years to design and tape out; the models it is built for can change in months. Betting the silicon on an architecture is betting that the architecture holds still long enough for the chip to earn its cost back. Sometimes it does. Sometimes the ground moves first, and the most efficient chip in the world turns out to be efficient at the wrong thing.

The chokepoints it did not escape

And Sohu does not stand alone, however singular its design. Follow its own supply chain one link upstream and the concentration is the one this beat returns to again and again. The chip is made at TSMC, on an island whose fabs the entire advanced-computing economy already leans on. The high-bandwidth memory stacked beside the logic comes from the same tight handful of suppliers — SK Hynix, Samsung, Micron — whose output is already spoken for years in advance by Nvidia and everyone else in the queue. Etched has not stepped outside the chokepoints that define its industry. It has bought a position inside them. A more specialized chip is not a more independent one; it rides the same few fabs, the same few memory makers, the same single seismic island, as the GPUs it means to unseat.

How much rides on how little

Here is the wager, stated plainly. A GPU's inefficiency is also its insurance. It wastes silicon on flexibility precisely so that it does not care what you run on it — and if the field abandons the transformer tomorrow for some successor, a state-space model, a diffusion approach, an architecture not yet named, the GPU simply runs the new thing somewhat worse and keeps earning. Sohu carries no such insurance. Everything that makes it fast is the same thing that makes it brittle: the transformer is not a setting it can change but the shape of its body.

If attention gives way to something else, the most specialized chip in the data center becomes the most obsolete, and the $21 billion rests on a single unhedged proposition — that the architecture the world settled on in 2017 is not a phase but a destination. The industry has spent three years being startled by what these models can do. Etched has just cast, in four-nanometer silicon on a Taiwanese production line, a very large bet that it has run out of ways to be startled. The chip is a beautiful piece of engineering, and it cannot change its mind. Somewhere upstream, that is either the smartest thing anyone built this year or the most expensive — and the difference will be decided not in the fab but in the next architecture nobody has drawn yet.

References

  1. TechCrunch — Etched's valuation doubles to $21B in a month
  2. GlobeNewswire — Etched raises $700M at a $21B valuation and completes first customer delivery to Jane Street
  3. TechCrunch — AI chip startup Etched hits $10.3B valuation from big-name investors
  4. Data Center Dynamics — Etched closes $300m round, doubles valuation to $10.3bn
  5. Spheron — Etched Sohu vs Nvidia: transformer ASIC vs GPU
The Friday Brief

One email. Every Friday.

The week's machines, money, and people — in under five minutes.