Benchmarks

ByteDance is training a 10-trillion-parameter model. That number can't tell you whether it's any good.

The headline figure is a ceiling on a model that doesn't exist yet, measured on nothing, against a rival that won't disclose its own count. Here's what a parameter number can and can't buy you.

A tower marked with ByteDance branding, the company behind TikTok and Douyin.

Image: N509FZ / Wikimedia Commons (CC BY-SA 4.0)

Here is a number that travelled around the world this week without a single measurement attached to it: ten trillion. ByteDance, the company behind TikTok, is training an AI model with as many as ten trillion parameters, the Financial Times reported on August 7, citing three people familiar with the project. The number did what big round numbers do. It got reposted as a fact, framed as a run at the frontier, and stacked against the current largest Chinese model — Moonshot's Kimi K3, at 2.8 trillion — to produce the tidy line that ByteDance's model is "3.6 times larger."

Three-point-six times larger than what, exactly, and larger in a way that buys you what? Those are the two questions almost nobody asked before hitting repost, and they are the whole story. Because a parameter count is the one figure in AI you can publish without running a test, and it is very close to the least informative number the field produces about whether a model is any good. Let me walk through why, because the walk is the point — not "distrust the number" as a reflex, but knowing exactly which questions dissolve it.

The number is a ceiling on a thing that doesn't exist yet

Start with the footnote, because it is doing more work than the headline. The FT's reporting, read carefully, does not describe a finished model with ten trillion parameters. It describes a model in pretraining — the long, expensive first phase where a model learns from raw data — with ten trillion as a candidate upper limit, not a settled specification. Pretraining a model this size typically runs three to six months, and the final parameter count is decided along the way, not at the start. So the number in circulation is not a result. It is the top of a range, on a system that will not exist in any usable form for months, reported second-hand from anonymous sources.

That matters before anything technical, because it sets the category. This is not a benchmark. It is not even a released model with undisclosed benchmarks, which is the usual thing I complain about. It is an intention, sized. And the honest way to write about an intention is to say what it is evidence of — ambition, compute access, strategy — and to refuse to let it masquerade as evidence of capability, because there is not yet any capability to have evidence about. Nothing has been measured. There is nothing to measure yet.

Total is not active, and the gap is enormous

Now suppose the model ships at ten trillion parameters exactly. The comparison to Kimi K3 still doesn't mean what the "3.6 times" line implies, and here is the mechanism. Almost every model at this scale is a mixture-of-experts. Instead of one dense network where every parameter fires on every token, the model is split into many specialised sub-networks — "experts" — and a router picks a small handful of them for each token of text. Kimi K3 is the public example: 2.8 trillion parameters in total, but the router selects just 16 of its 896 experts per token, so only about 104 billion parameters actually do work on any given word.

Read that again, because it is the number that got lost. The figure that touches your query is the active parameter count, and for Kimi K3 that is roughly 104 billion — less than four percent of the 2.8 trillion on the label. The total count tells you how much the model can store; the active count tells you how much it computes per token, which is far closer to what you experience as capability and exactly what you pay for at inference time. ByteDance has disclosed no architecture at all. If its ten-trillion-parameter model is a sparse mixture-of-experts — and at that scale it almost has to be — the number that matters could be a small single-digit fraction of the headline. "10 trillion versus 2.8 trillion" may well be a comparison of two storage numbers whose real working numbers are unknown on one side and, on the other, about a hundred billion. The label is not the engine.

Bigger stopped predicting better years ago

Even the total count, taken at face value, is a weak predictor of quality, and this is not a hunch — it is the settled lesson of the last several years of model-building. There was a period when scaling parameters reliably scaled capability, and the field over-learned it. Then came the correction: given a fixed amount of compute, a smaller model trained on more, cleaner data routinely beats a larger model trained on less. Capability is set by a bundle — training compute, data quantity and quality, the training recipe, and the post-training that turns a raw model into a useful one — and parameter count is only one term in it, and not the dominant one. This is why models rumoured to be a fraction of the size of the largest systems have matched or beaten them, and why every serious lab stopped bragging about parameters somewhere around the point they realised the number embarrassed them as often as it flattered them.

So "3.6 times larger" does not translate to 3.6 times better, or 3.6 times anything you would care about. It translates to 3.6 times more parameters, which is a statement about storage and cost, not about reasoning, or coding, or truthfulness, or any of the things a model is actually judged on. The most you can say is that ByteDance intends to spend more — more silicon, more power, more months — than Moonshot did. That is real. It is also not the same sentence as "a better model," and the reposts quietly swapped one for the other.

A parameter count is the one number in AI you can publish without running a test — which is exactly why it predicts so little about the test. — On the difference between a plan and a result

Compared to what? Anthropic won't say

The frame everyone reached for is that ByteDance is chasing Anthropic's Mythos, the frontier system Chinese labs have struggled to match. Fine — but that framing has a hole in it you could drive a data centre through. Anthropic has never disclosed a parameter count for Mythos. Not an estimate, not a range, nothing. The company has said only that its Fable and Mythos tiers run on the same underlying model and differ in safety and distribution, and it has been pointed about the fact that the only meaningful comparison between two models is one run under identical conditions on the same tests.

So the celebrated match-up — ByteDance's ten trillion against the frontier — is a comparison in which one side's number does not exist. You cannot rank a model you can count against a model you can't, on the axis of counting, and learn anything. If Mythos is, say, a densely-run system with a fraction of the raw parameters but far more training compute and better data, then "ten trillion" tells you ByteDance built a bigger warehouse, not that it built a better factory. "Compared to what" is the first question I ask of any claim, and here the answer is: compared to a number the rival deliberately never published. That should end the ranking, not start it.

What the number actually tells you

None of this is a knock on ByteDance, and it is worth being precise about the difference between significant and impressive, because they are not the same and the distinction is the entire discipline. The report is significant. Training a model at this scale is understood to involve on the order of tens of thousands of top-end GPUs — the FT's reporting points at roughly 30,000 — and committing that much scarce, expensive, export-controlled compute is a genuine signal. It tells you ByteDance is still betting on raw scale at a moment when much of the field has been talking about a "data wall" and pivoting to efficiency and post-training. It tells you the company has the capital and the chips to make that bet. It tells you where ByteDance wants to stand.

What it does not tell you is whether the model is good, because "good" is a measured property and nothing has been measured. Here is the checklist the number can't fill in:

  • How many parameters are active per token — the figure that governs cost and much of the capability — which depends on an architecture ByteDance hasn't described.
  • How it scores on any benchmark, held-out or otherwise, since the model isn't trained and no evaluations exist.
  • How it compares to Mythos or anything else run under identical conditions, which is the only comparison that means anything.
  • Whether the training data was clean or contaminated with the tests it will later be graded on — the question I care about most, and one you can only answer after the fact.
  • Whether it will be released with open weights, which is the difference between a claim you can check and a claim you have to take on faith.

That last one is where I will give credit in advance. ByteDance has open-sourced serious models before, and if the ten-trillion-parameter system arrives with downloadable weights and evaluations that outsiders can re-run, then the number stops being a press line and starts being a testable claim. That is the honest path, and it is available. Kimi K3 took it; its 2.8 trillion parameters and its active-expert routing are public precisely because Moonshot let people check. Until ByteDance does the same, the correct thing to hold in your head is not "the biggest model yet" but "a very large training run, in progress, whose quality is unknown and unmeasured."

The field will get a real result eventually — a shipped model, ideally with weights and independent evaluations and, if I am allowed to dream, error bars. When it does, we can ask the questions that matter: active parameters, training compute, contamination, performance under identical conditions. Those answers will either justify the excitement or deflate it, and either way they will be worth more than the headline. Ten trillion is a number you can print before the model exists. That is the surest sign it is not yet telling you the thing you want to know.

References

  1. The Next Web — ByteDance is training a 10-trillion-parameter model to chase the frontier
  2. XenoSpectrum — ByteDance Trains AI With Up to 10 Trillion Parameters, but Raw Counts Can't Measure the Gap With Anthropic
  3. MLQ — ByteDance Is Training a 10 Trillion-Parameter AI Model, Financial Times Reports
  4. Crypto Briefing — ByteDance reportedly plans 10 trillion total-parameter model with 30,000 GPUs
  5. Technology.org — ByteDance Trains 10 Trillion-Parameter AI Model
The Friday Brief

One email. Every Friday.

The week's machines, money, and people — in under five minutes.