Hands-on

OpenAI just made its cheap model 80% cheaper. I spent a week trying to break it.

Luna dropped to a fifth of its price three weeks after launch, Terra shed 20%, and the flagship didn't move a cent. I rebuilt my actual week on the cheap tier to find the seam. It's exactly where you'd expect, and closer to the surface than the price makes it look.

A smartphone displaying the OpenAI logo held in front of a blurred screen

Image: Focal Foto / Wikimedia Commons (CC BY-SA 4.0)

For one week I moved every small, boring job I normally hand to a good model onto the cheapest one OpenAI sells. The transcript cleanups, the inbox triage, the first drafts of things I'd rewrite anyway, the little scripts that read a folder and tell me what's in it. I did this because on July 30 OpenAI cut the price of GPT-5.6 Luna by 80 percent, and I wanted to know what an 80 percent discount actually buys you when the thing on the other end is a language model and not a plane ticket.

By Thursday I had a clear answer, and it is the answer you would guess if you've been paying attention: the cheap model is genuinely good at the jobs where being wrong is cheap, and it is exactly as unreliable as it always was at the jobs where being wrong is expensive. The price changed. The seam did not move. What changed is how tempting it now is to run straight past the seam because the meter barely ticks.

What actually got cheaper

Start with the numbers, because they're the whole reason anyone is talking about this. GPT-5.6 shipped on July 9 in three tiers, named — with the branding restraint we've come to expect — Luna, Terra, and Sol. Luna is the small, fast one. Terra is the middle. Sol is the flagship you reach for when you actually need it to think. Three weeks in, OpenAI repriced the bottom two and left the top alone.

  • Luna dropped from $1.00 / $6.00 to $0.20 / $1.20 per million input/output tokens — an 80 percent cut on both sides.
  • Terra dropped from $2.50 / $15.00 to $2.00 / $12.00 — a flatter 20 percent.
  • Sol did not move. It's still $5.00 / $30.00 per million tokens.
  • In place of a Sol price cut, OpenAI introduced a 'Fast mode' that replaces the old Priority Processing: up to 2.5x faster responses at twice the standard rate, or $10 / $60 per million tokens.

Read that list again and notice the shape of it. The two models that compete on price got cheaper. The one model that competes on being the best did not. Instead, the flagship got a new way to spend more, not less. If you want Sol to answer faster, you now pay a premium for the speed. That is not a company under pressure at the top of its range. That is a company under pressure at the bottom of it, which is a different and more interesting thing.

OpenAI's stated reason is efficiency. The company says improvements across its models, inference infrastructure, serving software, and context handling let it pass savings down, and — in the flourish everyone quoted — that Sol itself did some of the optimization work, rewriting production kernels to cut serving costs by a reported 20 percent and running token-generation experiments that improved efficiency by a further 15 percent. Take that at face value if you like. I'd just note that efficiency gains are always available to be announced on the day you need a reason to cut prices, and rarely on the days you don't.

Why now, three weeks in

A price cut three weeks after launch is not a victory lap. You don't discount a product by four-fifths that fast because it's selling too well. You do it because the number next to it on the buyer's spreadsheet is smaller than yours, and the buyer has stopped being sentimental about whose logo is on the model.

The context here is a market that has quietly turned into a price war. CNBC reported earlier this summer that Chinese open-weight models had climbed to roughly 46 percent of US enterprise token usage on OpenRouter, the router that a lot of companies use to shop models by the token, and at times were running ahead of US-origin models outright. When the cheap, capable alternative is one config change away, 'our model is a bit better' stops being a pricing argument. The buyer will take 'good enough and a fifth the price' for anything that isn't mission-critical, and most tokens aren't mission-critical.

OpenAI's own pitch for the cut leans into this. The company says Luna now beats a rival frontier model — it named Fable 5 — on the Agents' Last Exam evaluation at a reported 99 percent lower cost per task. I have no way to check that number, and neither do you; it's a vendor claim on a vendor-chosen benchmark, and the phrase 'cost per task' is doing a lot of quiet work in it. But the fact that the pitch is now 'cheaper per task' rather than 'smarter' tells you where the fight moved. The frontier labs spent two years selling capability. This week one of them sold a discount.

The frontier labs spent two years selling capability. This week one of them sold a discount.

The week: where the cheap model is fine

Here's the part that matters if you actually use this stuff. I don't buy tokens by the million — I'm a person with a subscription, not a platform. But the apps I live in do buy tokens by the million, and when Luna gets 80 percent cheaper, the calculus behind every 'summarize this,' 'clean this up,' 'sort this inbox' button in every tool I use changes. So I rebuilt a week of my own small automations — the little agents I keep in a folder and am mildly embarrassed by — to run on Luna at the new price, and I watched where they held and where they frayed.

They held in more places than I expected, and all of them rhyme. Luna is excellent at the jobs that are really pattern-matching wearing a task's clothing. Cleaning up a messy meeting transcript into readable notes: fine, and fast enough that I stopped noticing the wait. Reading a folder of forty PDFs and telling me which three mention a specific clause: fine. Turning a wall of my own bullet points into a first-draft paragraph I was always going to rewrite: fine, and honestly a little better than fine. Triaging an inbox into 'reply today / reply this week / archive': fine, with the caveat below.

For this whole category, the 80 percent cut is not a nice-to-have; it's the difference between 'a feature I use when I remember' and 'a thing that just runs in the background because it costs nothing to leave on.' That's the real product of a price cut like this. It doesn't make the model smarter. It makes you stop rationing it. And for bulk, low-stakes, high-volume work, not rationing it is most of the value. The cheap model was already good enough for these. Now it's good enough and you'll forget you're paying for it.

The week: where it broke, at 4pm on Thursday

The frays showed up exactly where they always do — the moment two tasks had to touch, and the moment being wrong cost something. This is the seam. It is not a Luna problem or a GPT-5.6 problem; it's the same seam I've been writing about in every agent I've ever tested. The cheaper model just makes it more tempting to route work across the seam that should have gone one tier up.

My inbox triage agent is the clean example. On its own, Luna sorts mail into buckets well. But I'd wired it to also draft short replies to anything in the 'reply today' pile, and to skip anything that looked like it needed a decision I hadn't made yet. That second instruction — the judgment call about when to stop and ask me — is where it started guessing. On Thursday afternoon it confidently drafted a warm, well-written reply agreeing to a call I had specifically not agreed to, because the email was friendly and the pattern of a friendly email is a friendly reply. A more expensive model doesn't reliably fix this either. But it hesitates more, and in agent work hesitation is a feature. Luna is tuned to be fast and cheap, and fast and cheap reads, in practice, as decisive. A decisive guess wearing a sent-mail timestamp is just a mistake you have to go apologize for.

The other failure was quieter and more instructive. I gave the same multi-step research task — read these sources, reconcile where they disagree, tell me what's actually known — to Luna and to Sol, back to back. Luna produced a fluent, confident, tidy answer that flattened the disagreements out of existence. It read beautifully and it was subtly wrong in the way that costs you most, because nothing in the prose warns you. Sol was slower, hedged more, and told me which two sources conflicted and why. That difference is the entire ballgame, and it's the one thing an 80 percent discount cannot buy you. You are not paying Sol for better sentences. You're paying it to know when not to be sure.

A decisive guess wearing a sent-mail timestamp is just a mistake you have to go apologize for.

The Fast-mode tell

Come back to the thing OpenAI didn't cut. Sol's price held, and the only new option at the top of the range is a way to pay double for speed. I think this is the most honest sentence in the whole announcement, and OpenAI didn't write it in words — it wrote it in the price list.

What it's conceding is that the two ends of the range are now different businesses. At the bottom, models are a commodity, sold by the token against Chinese open weights and everyone else, and the only lever is price, so the price fell through the floor. At the top, the flagship still has something people will pay for — judgment, the hesitation, the not-being-wrong-expensively — and there the company isn't discounting at all; it's finding new things to charge for. The cut and the Fast mode are the same strategy read from both ends: give the cheap stuff away to keep the volume, and protect the margin where the model is genuinely hard to replace.

For a buyer, that's clarifying. It means the pricing is finally telling you the truth the benchmarks obscure. Use the cheap tier for anything a commodity can do, which is more than the labs would like to admit. Pay up only for the work where being sure matters, which is less than you think but never zero. The price sheet has quietly sorted your own tasks into those two piles for you.

So who should switch

If you build or run anything on the API, the answer is easy: move your bulk, low-stakes calls to Luna today, keep a hard rule that anything touching a decision, a customer, or a fact you can't afford to get wrong escalates to Sol, and don't let the tiny meter tempt you into blurring that line. The 80 percent isn't a reason to use the cheap model for more kinds of work. It's a reason to use it for the same work, far more often, and to feel the savings in your bill instead of in your error rate.

If you're a normal person with a ChatGPT subscription, none of this changes your Tuesday directly — you don't see token prices, and the app picks models for you. But it changes the ground under every tool you pay for, because the thing those tools spend on your behalf just got a lot cheaper at the low end, and that will show up as features that used to be metered quietly becoming free. That's the good version. The bad version is the one I hit on Thursday: cheaper models make it cheaper to let software make decisions it shouldn't, and the invitation to do that is now stronger than it was in June.

The scorecard, since I keep one: over the week I moved four of my little agents permanently onto Luna and they're better off there, mostly because I've stopped thinking about what they cost. I moved one — the one that drafts replies — back to Sol and put a hard stop in front of anything it wants to send, because the cheap version's confidence is a liability exactly proportional to how good its prose is. Four kept, one walked back. The price cut is real, the savings are real, and the seam is precisely where it was three weeks ago, now with a smaller number in front of it and a stronger reason to pretend it isn't there.

References

  1. OpenAI cuts prices for two of its GPT-5.6 AI models as companies grow sensitive to costs — CNBC
  2. OpenAI cuts GPT-5.6 Luna pricing by 80% and accelerates Sol API — Crypto Briefing
  3. AI price wars: OpenAI cuts GPT-5.6 Luna prices by 80% as model competition shifts toward cost — VentureBeat
  4. OpenAI Cuts GPT-5.6 Pricing Up To 80%, As AI Costs Come Under Scrutiny — Forbes
  5. GPT-5.6: Frontier intelligence that scales with your ambition — OpenAI
The Friday Brief

One email. Every Friday.

The week's machines, money, and people — in under five minutes.