GPT 5.6 Pricing Just Fell 80% on Luna. Price Isn't the Cost.
GPT 5.6 pricing changed on July 30, 2026. Luna dropped 80%, to $0.20 per million input tokens and $1.20 per million output tokens. Terra dropped 20%, to $2 and $12. Sol, the model most people run for hard work, kept its rate, gaining only an API side Fast mode that trades twice the price for up to 2.5 times the speed.
Here is the part the headline buries. A rate is one half of a multiplication, and the other half is how many tokens the model decides to spend. You control that half, with a setting you already have, and on OpenAI's own chart it swings the cost of a task about fivefold on a single model. That is a wider range than the distance between several of the separately priced models sitting next to it, and unlike a price cut you do not have to wait for anyone to hand it to you.
Key Takeaways
- Luna got the 80% cut, Terra 20%, and Sol's rate held with a faster paid tier added on top.
- A rate is not a cost. Cost is the rate multiplied by tokens spent, and the model largely chooses the second number.
- The effort dial moves cost per task about fivefold on a single model, per OpenAI's own chart. Try raising effort on a cheap model before switching to an expensive one.
- OpenAI's own advice with this release is to stop defaulting to its most expensive model for work that does not need it.
- Match the tier to whether a task has a checkable answer, not to how important the task feels.
- In Codex and ChatGPT Work, OpenAI says Luna and Terra usage now counts for less against your allowance, so routing work down is how you collect this cut without an API key.
- Building on the API? Every response hands back the exact token count it charged you, thinking included.
Short on time? Skip straight to the 4 things worth changing this week. The rest of this explains why they work.
What actually changed in GPT 5.6 pricing
At the time of writing, this is the shape of the change, per the announcement.
| Model | New price per million tokens | Change |
|---|---|---|
| Luna | $0.20 in, $1.20 out | Down 80% |
| Terra | $2 in, $12 out | Down 20% |
| Sol | Rate unchanged | New Fast mode, up to 2.5x quicker at twice the price |
Two things are worth pulling out of that table. The cut is concentrated in the smallest model, and the model doing the heavy work got no discount at all. Sol's Fast mode is an API only option and most people on a subscription can ignore it, but it is worth one line for what it prices: double the rate for speed, with, in OpenAI's own words, "no change in intelligence."
The chart published with the announcement is worth reading closely, because of what it puts on the horizontal axis.

That axis is cost per task, not price per token, which is the honest way to measure this and not the number in the headline. The vertical axis is the Artificial Analysis Intelligence Index, a third party score rather than an in house benchmark, which is worth a little more trust than a vendor grading its own work. OpenAI still chose which models to plot against it.
Then look at the blue line. Luna is not a dot on that chart, it is a curve, and it travels roughly fivefold along the cost axis while climbing about eighteen points of score. One model, one price per token, a fivefold spread in what a task actually costs. That spread is wider than the distance between several of the separate, differently priced models plotted beside it.
Now notice what is missing, because it is the most instructive thing on the page. Luna is the only model drawn as a curve. Every competitor is a single dot. Read their labels: "Claude Opus 5 Low", "Claude Sonnet 5 High", "GLM-5.2 Max", "Claude Haiku Reasoning". Each of those names one effort setting out of several that model has, and the other settings are not plotted. Opus 5 has a curve of its own and you cannot see it here.
That is ordinary practice. You plot your own model in depth because you have the data, and everyone else at a representative point, and every vendor's chart does some version of it. It is worth noticing anyway, because it tells you what the chart can and cannot answer. You get Luna's full cost quality curve, and one sampled setting from everyone else. So a comparison across models here is a curve against a snapshot, which is a fine thing to be curious about and a thin thing to conclude from.
The honest takeaway is the one about Luna alone, and it survives all of that: moving the effort dial on a single model changes cost per task by about fivefold. That is a wider swing than the distance between several of the separately priced models on the same plot, and it costs nothing to try.
Everything below follows from that.
For scale, Anthropic's pricing docs list Claude Opus 5 at $5 per million input tokens and $25 per million output tokens. Set against Luna's new $0.20 and $1.20, that is 25 times on input and about 21 times on output. Output is the number to watch, because a model's own thinking is charged at the output rate, so the harder it thinks the more it draws on the expensive side.
That gap still sounds decisive, and it is only half a calculation. The other half is how many tokens each model spends to finish the same job, which no price sheet can tell you.
One more caveat belongs on that chart, and it cuts against my own reading of it. Cached input reads for roughly a tenth of fresh input on both major providers, and a long agentic session runs warm, so a large share of real input never bills at the list rate. A cost per task figure computed on an eval suite has no reason to model that, and I have not seen one that says it does. If that is right, the axis describes a world without caching, and an input to output mix nothing like a coding session where the same context is re read hundreds of times. It is still the right axis to reason on. It is not your number.
Why a cheaper model can cost you more
First, who pays what, because the numbers above are API prices and most people reading this never see one.
If you build on the API, a rate cut is money, directly. If you work inside a subscription like Codex or Claude Code, you pay the same amount every month no matter what you run. There is no invoice to watch. What you have instead is an allowance that drains as you work, and when it runs out your afternoon stops. That is the real currency, and this whole post is about spending it more slowly.
Both cases obey the same arithmetic, so hold on to this one line: what you spend is the rate multiplied by the number of tokens the model decides to use. A reasoning model chooses how long to think before it answers, and that thinking counts. The rate is fixed and published. The second number is not, it swings enormously, and it is the one nobody puts in a headline.

That is the whole idea. A single token is almost free and you are buying millions of them, so what a session costs you is set by the size of the block, not the price of the square. On the API that shows up as money. On a subscription it shows up as how much of your day you got.
There is one piece of genuinely good news buried here for subscription users, and it went almost unmentioned in the coverage. In its own posts announcing this change, OpenAI said the lower Luna and Terra prices are reflected in how usage is counted in Codex and ChatGPT Work, so the same allowance now goes further on those two models. Read the scope carefully, because it matters: the company named Codex and ChatGPT Work, not Plus or Pro. If you are in Codex, this is the rare price change that reaches you without an API key, and the way to collect it is specific. It only applies to Luna and Terra. Work you leave on Sol collects nothing.
Which turns the abstract advice into something concrete. For a Codex user, routing bounded work down to Terra or Luna is not merely tidier, it is now the mechanism by which this announcement makes your week longer.
This is where my own experience stops matching the headline. I run Claude Opus 5 as my main driver, with Sonnet alongside it because the compute allows it, and I use both GPT 5.6 Terra and Sol, Sol as an orchestrator in short bursts and Terra for a good deal of everyday work. Sol is genuinely good. It is also the top tier of that family and a heavy thinker, which is an expensive combination: the highest rate applied to a lot of tokens. It reasons at length, and length is the thing you are buying. A price cut on a different model in the same family does nothing about that.
So a rate cut is real money, and it is one of two levers rather than the only one. The other is how much thinking you ask for.
Which of the two matters more is not fixed, and it is worth being precise about, because the easy version of this argument is wrong in both directions.
Between models priced close together, effort dominates. Push the cheaper one to a high setting and you can easily spend past the pricier one running at a low setting. It is an easy trap, and I have walked into it.
Across a wide gap the rate matters more, and it matters in favour of the cheap model. Luna and Opus 5 are twenty five times apart on input and about twenty one times apart on output. On the one comparison the chart lets us make, Luna at its most expensive setting still comes out cheaper per task than the plotted point for Opus 5 at its cheapest, which makes the cheap model working hard a real option rather than a compromise. Note the size of that claim: it rests on a single plotted point for Opus 5, and the rest of that model's curve is not shown.
Now put the rate card next to the chart, because they do not say the same thing. On rates, Luna is twenty five times cheaper. On the chart's cost per task axis, comparing Luna at its top setting against the plotted point for Opus 5 at low, the gap is closer to five times. Still a large advantage. Five times, not twenty five.
Be careful about what that does and does not prove, because the obvious conclusion is not available here. It is tempting to say the difference is token efficiency, that Opus 5 gets there in fewer tokens. Maybe it does. You cannot tell from this chart, because the two points being compared sit at opposite ends of their own effort ranges: Luna is at its most expensive setting and Opus 5 at its cheapest, and neither model's other points are drawn. Luna's own curve already spans about fivefold on its own, which is the entire size of the remaining gap, so the compression could come from the dial alone.
That uncertainty is the finding, not a hole in it. A setting you can change in a second moves cost per task by roughly as much as the entire remaining distance between a cheap model and a frontier one. The two levers are in the same weight class, though only one of them gets a press release. Whatever a twenty percent price cut just handed you, the dial in front of you is playing for comparable stakes, and it is already yours.
Sol is the case that broke my tidy rule. In my own use it spends faster than the others even with the effort turned down, which is why I keep it as an orchestrator in short bursts rather than as the thing running all day.
That points at something worth knowing generally, and it is the part I got wrong for a while. Effort settings are not calibrated across models. Low, medium and high are each vendor's own labels for their own model, not units on a shared scale, so medium on one model is not medium on another and there is no published conversion between them. That is exactly why the equivalences people trade in forum threads keep sounding authoritative and keep aging badly. When you move a workload to a new model, the setting you were using does not carry over as a number. You are calibrating from scratch, and the only way to find the new level is on your own work.
The vendor is telling you to route down
The most useful sentence in the announcement is not about price at all. Describing Terra, its middle model, OpenAI writes:
If you currently default to Sol for work that doesn't require maximum intelligence, consider evaluating Terra.
Read that again with the incentives in mind. The company is telling paying customers to move work off its most expensive model, in the same release that left that model's price where it was.
I would not read that purely as candour, though. Frontier models are generally the capacity constrained ones, and steering everyday work onto a mid tier model is good business: it frees the scarce thing, and depending on the economics it can help margin too. Guidance of this shape is common across large providers, and it is reasonable to assume it serves them as well as it serves you.
The useful part survives the cynicism, though. A vendor that thought its top model was being used correctly would have no reason to write that sentence at all. It reads as aimed at a habit the company is well placed to observe, which is people reaching for the strongest available model by default. You do not need to accept the motive to accept the observation.
It also sets up the harder question, which is how you decide what "does not require maximum intelligence" means for your own work. The announcement does not answer that. Nothing can, except a test.
The reaction told the real story
The threads that followed the announcement are more useful than the announcement.
The first thing the community did was deflate the number. The top comment on the largest thread pointed out that only Luna got the 5x reduction and that Sol had not moved. The second corrected the headline outright. Nobody led with celebration.
The second thing they did is the part worth sitting with. They asked for a usage reset. An entire thread turned into people requesting more capacity, hours after being handed cheaper tokens. That is the whole argument in one gesture. The constraint people actually feel is how fast a session consumes their window.
Buried under all of it was the genuinely useful material, sitting in replies where almost nobody will find it. People were trading equivalences. One reported Luna at its highest reasoning setting landing close to Terra at medium. Another put Luna at maximum near Opus 5 at medium. A third described using the big model to plan, a small one to implement, and a small one again to review its own direction. Someone else found Luna good enough for 90% of a data extraction job and switched.
Every one of those is a routing decision, and every one of them is folklore. It is unsourced, it is scattered, and it is pinned to model names that will be wrong within a couple of months. The useful part is what produced them. Every one of those people ran a test on their own work.
How I pick an effort level
I do not keep a default. I start at the level I think the job needs, and often one notch below that to see whether the cheaper setting holds. Then I read the output and adjust.
That last step is the whole method, and it is why a fixed rule decays. If I decide today that medium is right for a given kind of work, that decision is attached to this month's version of this model. The next release changes what each level actually buys, and my setting quietly becomes wrong without ever announcing itself.
A concrete one, and I should be straight about how loosely I got there. I expected video editing work to need a large model at a high setting. It runs fine on Terra at medium. That came from noticing it across a stretch of work, not from the protocol below, so treat it as the sort of hunch worth testing rather than a result I can hand you. I would not have stumbled on it by reading a benchmark, and I would not have found it while sitting on a default I set weeks earlier.
The pattern generalises past any one vendor. Being wrong about a setting is expensive in a quiet way, because you pay the difference on every task afterwards until you happen to notice.
A word on what testing actually takes, because I had a looser version of this advice in an earlier draft and it does not hold up. Running a task once at each setting and going with whichever output you preferred feels like a test, and I have done it plenty of times. It is not one. These models are stochastic, so a single run tells you about that run, and it is precisely how the equivalences in those threads got made.
The cheap version that does hold up has three parts. Decide what "it worked" means before you run anything, in one written line, because deciding afterwards means grading the output against the output. Run each setting more than once, since two passes catch the case where one run got lucky. And pick tasks where you can tell pass from fail without arguing with yourself, which usually means something with a test suite, a checkable list, or an obvious right answer.
Then you need to measure what it cost, and how you do that depends on which side you are on.
On a subscription, your meter is the allowance. You have no per token bill to read, so what you watch is how far your window moves. That reading is coarse, and a single small task may not shift it at all, so measure across a batch: note where you are, run the whole set on one configuration, note where you ended up, then start fresh and do the same on the other. Two numbers for the same six tasks tells you plenty.
On the API, the exact number is handed to you. Every response comes back with a count of the tokens it just charged for. OpenAI breaks out how many were reasoning and how many came from cache instead of being read fresh; Anthropic reports cache reads and writes separately, though it folds thinking into the output number. That is a precise measurement, already paid for, sitting in the reply. Nothing needs estimating, which is worth knowing given how much of this conversation runs on guesswork.
Either way, read it on every run rather than on one, because these counts wander. The amount of thinking is itself a choice the model makes fresh each time, so one reading gives you a single sample of a moving number. You are already repeating the runs for quality. The cost data rides along free, and averaging it is what turns an anecdote into a measurement.
That is an afternoon, not a research project. And the answer comes from your own work, which is the only place it counts.
If you want the Claude side of this in more detail, I covered effort levels and where the cheaper model already wins, and separately the case for deleting instructions that made older models reliable and now just cost tokens.
4 things worth changing this week
Concrete, in the order I would do them.
1. Raise effort on a small model before you reach for a big one. This is the one that follows directly from that curve, and it runs against the usual reflex. When output is not good enough, the instinct is to switch to the stronger model. Try turning the effort up on the model you are already using first. On the chart above, Luna's own range covers about eighteen points of index score, which is a wider band than the gap between plenty of separately priced models. A cheap model thinking harder is often cheaper than an expensive model thinking casually, and it is the least disruptive thing on this list to check, since you change a setting rather than your whole setup.
2. Match the tier to the shape of the work, not to the importance of the work. The instinct is to give important things the best model. Importance is not the variable. What matters is whether the task has one right answer you can check.
| Shape of the task | Where I start |
|---|---|
| Bounded and checkable: a defined change, a test to pass, a format to fill | Cheapest tier, effort turned up |
| High volume and repetitive: extraction, classification, summarising, triage | Cheapest tier, effort low, and run it through a batch endpoint |
| Ambiguous, or the plan matters more than the code | Strongest tier, short burst, then hand the plan down |
| Long and open ended, where it has to hold a lot at once | Strong tier, and watch the window |
That third row is where the money goes, and it is usually the smallest share of the work. Planning is the part worth paying for, since a bad plan executed cheaply is still a bad plan.
On the second row, one more thing that has been sitting there the whole time and dwarfs this week's news: if the work does not need an answer back immediately, batch endpoints run around half price at both major providers. That is a larger discount than Terra just received, and nobody announced it today. Like Fast mode, this one lives on the API, so skip it if you work inside a subscription.
3. Know what a model switch actually costs you, because it is smaller than it feels. That cache discount from earlier is per model, so moving a running session across means re reading the whole context at full price on the other side. On Anthropic, where you mark what to cache, priming it costs more than plain input as well. On OpenAI, caching is automatic and there is no write premium.
It is tempting to conclude you should never switch mid session. That is the wrong lesson, and I had it backwards until I did the arithmetic. The cache cost is paid once. The rate difference is paid on every turn that follows. So the longer the session still has to run, the more the switch pays for itself. It only loses when the session is nearly over, or when you keep flip flopping and pay the re prime again and again. Decide by how much work is left, not by how much you have already put in. On a subscription the same thing plays out in allowance rather than dollars: you re read the context either way, and it comes out of your window.
4. If you are on a subscription, stop thinking in dollars and go collect the discount. Your price is fixed, so the rate cut never shows up on a bill. Your scarce resource is the window, and everything above spends it: a heavier model, a higher effort setting, a longer context, a cold cache. For Codex and ChatGPT Work specifically, OpenAI has said the lower Luna and Terra prices are reflected in how usage is counted, so moving bounded work onto those two models is now the way this announcement reaches you at all. Leave that work on Sol and you collect none of it.
The chart changed what I am running next
Writing this changed my own plan, which is not how I expected it to go.
Luna is a model I almost never touch. My daily driver is Claude Opus 5 at low effort. Look at where those two sit on the chart above: Luna's top point lands at roughly the same index score as the single plotted point for Opus 5 at low, at something like a fifth of the cost per task. I was not looking for that. I was looking at a price announcement about a model I do not use.
So I am going to run Luna against my own work, on the protocol above, and find out.
I want to be exact about what that is and is not, and the caveat from earlier applies to me as much as to anyone. That comparison is a full curve against one sampled dot, chosen by OpenAI, on one third party index, measured on tasks that are not mine. It is not evidence that Luna matches Opus 5. It is a hypothesis with a cheap experiment attached, which is the only thing a chart like that can ever honestly give you.
The result I would find most useful is the negative one, where it fails on something Opus 5 handles and I go back. That is the easiest part of the loop to skip, and it is the only part that produces anything you can trust.
If you are new to this, here is the part that matters
Everything above assumes you already run these tools daily. If you do not, most of it reads as noise about products you have never opened. Two things survive the jargon.
What you are buying is attention on a task, and you can ask for more or less of it. I spent a long time treating the choice of model as the decision and the amount of thinking as a detail. They are closer to equal partners, and only one of them is free to change.
And this is a skill rather than a fact to memorise. Prices moved this week and will move again, so anyone handing you a definitive table of which model to use for what is describing a snapshot that is already going stale. Knowing how to check on your own work costs an afternoon and keeps paying after every release.
That is also most of what I teach in the AI workshop, if you want to build the habit with someone watching rather than by burning your own limit on it.
