GPT 5.6 Pricing Just Fell 80% on Luna. Price Isn't the Cost.
GPT 5.6 pricing changed on July 30, 2026. Luna dropped 80%, to $0.20 per million input tokens and $1.20 per million output tokens. Terra dropped 20%, to $2 and $12. Sol, the model most people run for hard work, kept its rate, gaining only an API side Fast mode that trades twice the price for up to 2.5 times the speed.
Here is the part the headline buries. A rate is one half of a multiplication, and the other half is how many tokens the model decides to spend. You control that half, with a setting you already have, and on OpenAI's own chart it swings the cost of a task about fivefold on a single model. That is a wider range than the distance between several of the separately priced models sitting next to it, and unlike a price cut you do not have to wait for anyone to hand it to you.
Key Takeaways
- Luna got the 80% cut, Terra 20%, and Sol's rate held with a faster paid tier added on top.
- A rate is not a cost. Cost is the rate multiplied by tokens spent, and the model largely chooses the second number.
- The effort dial moves cost per task about fivefold on a single model, per OpenAI's own chart. Try raising effort on a cheap model before switching to an expensive one.
- OpenAI's own advice with this release is to stop defaulting to its most expensive model for work that does not need it.
- Match the tier to whether a task has a checkable answer, not to how important the task feels.
- In Codex and ChatGPT Work, OpenAI says Luna and Terra usage now counts for less against your allowance, so routing work down is how you collect this cut without an API key.
- Building on the API? Every response hands back the exact token count it charged you, thinking included.
Short on time? Skip straight to the 4 things worth changing this week. The rest of this explains why they work.
What actually changed in GPT 5.6 pricing
At the time of writing, this is the shape of the change, per the announcement.
| Model | New price per million tokens | Change |
|---|---|---|
| Luna | $0.20 in, $1.20 out | Down 80% |
| Terra | $2 in, $12 out | Down 20% |
| Sol | Rate unchanged | New Fast mode, up to 2.5x quicker at twice the price |
Two things are worth pulling out of that table. The cut is concentrated in the smallest model, and the model doing the heavy work got no discount at all. All three carry the same context ceiling, and how much of it Codex actually hands you is a setting rather than a property of the model you pick. Sol's Fast mode is an API only option and most people on a subscription can ignore it, but it is worth one line for what it prices: double the rate for speed, with, in OpenAI's own words, "no change in intelligence."
The chart published with the announcement is worth reading closely, because of what it puts on the horizontal axis.

That axis is cost per task, not price per token, which is the honest way to measure this and not the number in the headline. The vertical axis is the Artificial Analysis Intelligence Index, a third party score rather than an in house benchmark, which is worth a little more trust than a vendor grading its own work. OpenAI still chose which models to plot against it.
Then look at the blue line. Luna is not a dot on that chart, it is a curve, and it travels roughly fivefold along the cost axis while climbing about eighteen points of score. One model, one price per token, a fivefold spread in what a task actually costs. That spread is wider than the distance between several of the separate, differently priced models plotted beside it.
Now notice what is missing, because it is the most instructive thing on the page. Luna is the only model drawn as a curve. Every competitor is a single dot. Read their labels: "Claude Opus 5 Low", "Claude Sonnet 5 High", "GLM-5.2 Max", "Claude Haiku Reasoning". Each of those names one effort setting out of several that model has, and the other settings are not plotted. Opus 5 has a curve of its own and you cannot see it here.
That is ordinary practice. You plot your own model in depth because you have the data, and everyone else at a representative point, and every vendor's chart does some version of it. It is worth noticing anyway, because it tells you what the chart can and cannot answer. You get Luna's full cost quality curve, and one sampled setting from everyone else. So a comparison across models here is a curve against a snapshot, which is a fine thing to be curious about and a thin thing to conclude from.
The honest takeaway is the one about Luna alone, and it survives all of that: moving the effort dial on a single model changes cost per task by about fivefold. That is a wider swing than the distance between several of the separately priced models on the same plot, and it costs nothing to try.
Everything below follows from that.
For scale, Anthropic's pricing docs list Claude Opus 5 at $5 per million input tokens and $25 per million output tokens. Set against Luna's new $0.20 and $1.20, that is 25 times on input and about 21 times on output. Output is the number to watch, because a model's own thinking is charged at the output rate, so the harder it thinks the more it draws on the expensive side.
That gap still sounds decisive, and it is only half a calculation. The other half is how many tokens each model spends to finish the same job, which no price sheet can tell you.
One more caveat belongs on that chart, and it cuts against my own reading of it. Cached input reads for roughly a tenth of fresh input on both major providers, and a long agentic session runs warm, so a large share of real input never bills at the list rate. A cost per task figure computed on an eval suite has no reason to model that, and I have not seen one that says it does. If that is right, the axis describes a world without caching, and an input to output mix nothing like a coding session where the same context is re read hundreds of times. It is still the right axis to reason on. It is not your number.
Why a cheaper model can cost you more
First, who pays what, because the numbers above are API prices and most people reading this never see one.
If you build on the API, a rate cut is money, directly. If you work inside a subscription like Codex or Claude Code, you pay the same amount every month no matter what you run. There is no invoice to watch. What you have instead is an allowance that drains as you work, and when it runs out your afternoon stops. That is the real currency, and this whole post is about spending it more slowly.
Both cases obey the same arithmetic, so hold on to this one line: what you spend is the rate multiplied by the number of tokens the model decides to use. A reasoning model chooses how long to think before it answers, and that thinking counts. The rate is fixed and published. The second number is not, it swings enormously, and it is the one nobody puts in a headline.

That is the whole idea. A single token is almost free and you are buying millions of them, so what a session costs you is set by the size of the block, not the price of the square. On the API that shows up as money. On a subscription it shows up as how much of your day you got.
There is one piece of genuinely good news buried here for subscription users, and it went almost unmentioned in the coverage. In its own posts announcing this change, OpenAI said the lower Luna and Terra prices are reflected in how usage is counted in Codex and ChatGPT Work, so the same allowance now goes further on those two models. Read the scope carefully, because it matters: the company named Codex and ChatGPT Work, not Plus or Pro. If you are in Codex, this is the rare price change that reaches you without an API key, and the way to collect it is specific. It only applies to Luna and Terra. Work you leave on Sol collects nothing.
Which turns the abstract advice into something concrete. For a Codex user, routing bounded work down to Terra or Luna is not merely tidier, it is now the mechanism by which this announcement makes your week longer.
This is where my own experience stops matching the headline. I run Claude Opus 5 as my main driver, with Sonnet alongside it because the compute allows it, and I use both GPT 5.6 Terra and Sol, Sol as an orchestrator in short bursts and Terra for a good deal of everyday work. Sol is genuinely good. It is also the top tier of that family and a heavy thinker, which is an expensive combination: the highest rate applied to a lot of tokens. It reasons at length, and length is the thing you are buying. A price cut on a different model in the same family does nothing about that.
So a rate cut is real money, and it is one of two levers rather than the only one. The other is how much thinking you ask for.
Which of the two matters more is not fixed, and it is worth being precise about, because the easy version of this argument is wrong in both directions.
Between models priced close together, effort dominates. Push the cheaper one to a high setting and you can easily spend past the pricier one running at a low setting. It is an easy trap, and I have walked into it.
Across a wide gap the rate matters more, and it matters in favour of the cheap model. Luna and Opus 5 are twenty five times apart on input and about twenty one times apart on output. On the one comparison the chart lets us make, Luna at its most expensive setting still comes out cheaper per task than the plotted point for Opus 5 at its cheapest, which makes the cheap model working hard a real option rather than a compromise. Note the size of that claim: it rests on a single plotted point for Opus 5, and the rest of that model's curve is not shown.
Now put the rate card next to the chart, because they do not say the same thing. On rates, Luna is twenty five times cheaper. On the chart's cost per task axis, comparing Luna at its top setting against the plotted point for Opus 5 at low, the gap is closer to five times. Still a large advantage. Five times, not twenty five.
Be careful about what that does and does not prove, because the obvious conclusion is not available here. It is tempting to say the difference is token efficiency, that Opus 5 gets there in fewer tokens. Maybe it does. You cannot tell from this chart, because the two points being compared sit at opposite ends of their own effort ranges: Luna is at its most expensive setting and Opus 5 at its cheapest, and neither model's other points are drawn. Luna's own curve already spans about fivefold on its own, which is the entire size of the remaining gap, so the compression could come from the dial alone.
That uncertainty is the finding, not a hole in it. A setting you can change in a second moves cost per task by roughly as much as the entire remaining distance between a cheap model and a frontier one. The two levers are in the same weight class, though only one of them gets a press release. Whatever a twenty percent price cut just handed you, the dial in front of you is playing for comparable stakes, and it is already yours.
Sol is the case that broke my tidy rule. In my own use it spends faster than the others even with the effort turned down, which is why I keep it as an orchestrator in short bursts rather than as the thing running all day.
That points at something worth knowing generally, and it is the part I got wrong for a while. Effort settings are not calibrated across models. Low, medium and high are each vendor's own labels for their own model, not units on a shared scale, so medium on one model is not medium on another and there is no published conversion between them. That is exactly why the equivalences people trade in forum threads keep sounding authoritative and keep aging badly. When you move a workload to a new model, the setting you were using does not carry over as a number. You are calibrating from scratch, and the only way to find the new level is on your own work.
The vendor is telling you to route down
The most useful sentence in the announcement is not about price at all. Describing Terra, its middle model, OpenAI writes:
If you currently default to Sol for work that doesn't require maximum intelligence, consider evaluating Terra.
Read that again with the incentives in mind. The company is telling paying customers to move work off its most expensive model, in the same release that left that model's price where it was.
I would not read that purely as candour, though. Frontier models are generally the capacity constrained ones, and steering everyday work onto a mid tier model is good business: it frees the scarce thing. Guidance of this shape is common across large providers, and it is reasonable to assume it serves them as well as it serves you.
The useful part survives the cynicism, though. A vendor that thought its top model was being used correctly would have no reason to write that sentence at all. It reads as aimed at a habit the company is well placed to observe, which is people reaching for the strongest available model by default. You do not need to accept the motive to accept the observation.
It also sets up the harder question, which is how you decide what "does not require maximum intelligence" means for your own work. The announcement does not answer that. Nothing can, except a test.
