Claude Opus 5 vs Fable 5: 8 Wins, 4 Losses, Half Cost
Claude Opus 5 is Anthropic's new Opus model, released on July 24, 2026, and on Anthropic's own comparison table it beats Claude Fable 5 on 8 of 12 benchmarks at half Fable 5's price. It costs the same as the model it replaces, Opus 4.8, at $5 per million input tokens and $25 per million output tokens at the time of writing. It does not win everything. Fable 5 still takes four rows, two of them coding. This post walks through the full scoreboard, where each model actually wins, what the charts leave out, which effort setting I run it on, and how it has behaved for me since launch day.
The AI space moves weekly, so every price and benchmark below is current at the time of writing and linked to its source.
Key Takeaways
- Claude Opus 5 wins 8 of the 12 head to head benchmark rows against Claude Fable 5 and loses 4, two of which are coding rows.
- Three of those four losses are under a single point, so the honest read is that the two models are close on the rows Fable 5 takes.
- At the time of writing Opus 5 costs $5 per million input tokens and $25 per million output tokens, the same as Opus 4.8 and half of Fable 5.
- One of the eight wins is uncontested, because Fable 5 posted no score on novel problem solving, so the stricter count is seven clear wins plus one walkover.
- Effort matters as much as the model choice. On Anthropic's charts, Opus 5 at low effort sits above Opus 4.8 at its most expensive setting, which is why low is where I start and the task decides whether it moves up.
- Opus 5 is not the top model on every chart. GPT-5.6 Sol sits at or above it on the Artificial Analysis coding index across most of the price range.
What Is Claude Opus 5?
Claude Opus 5 is the newest model in Anthropic's Opus line, built for agentic coding, long running tasks, and knowledge work. Anthropic released it on July 24, 2026 as the new default on the Claude Max plan and the strongest model available on Claude Pro. For anyone on a Claude subscription rather than paying per token, that is the part that matters: the model behind your daily work changed, and you did not have to do anything.
The pricing did not move. Anthropic states plainly that Opus 5 "provides greatly improved performance for the same cost as its predecessor, Opus 4.8."
At the time of writing that means $5 per million input tokens and $25 per million output tokens, per Anthropic's announcement. Claude Fable 5, the model Opus 5 is measured against throughout this post, sits at exactly double that, $10 and $50 per million tokens.
Cost per task is a different measure from price per token, and the two do not move together. A model that thinks longer can cost more per task even at the same token price, which is why the charts further down plot cost per task rather than sticker price.
Claude Opus 5 vs Claude Fable 5: the Full Scoreboard
Claude Opus 5 wins 8 of the 12 head to head rows on Anthropic's own comparison table, and Claude Fable 5 wins 4. Here is the whole comparison, including the rows Opus 5 loses.
One thing to know about the source before reading it. Anthropic's table has four columns, not two: Opus 5, Fable 5, Opus 4.8, and GPT-5.6 Sol. What follows is the Opus 5 against Fable 5 slice of it, which is the comparison most people want, and it is not the same as asking which model tops each row outright.
| Benchmark (what it measures) | Opus 5 | Fable 5 |
|---|---|---|
| Frontier-Bench v0.1, agentic terminal coding | 43.3 | 33.7 |
| GDPval-AA v2, knowledge work | 1861 | 1747 |
| ARC-AGI-3, novel problem solving | 30.2 | no score |
| BrowseComp, agentic search | 90.8 | 87.4 |
| Humanity's Last Exam, reasoning with tools | 64.7 | 63.9 |
| OSWorld 2.0, computer use | 70.6 | 66.1 |
| AutomationBench, business workflows | 26.0 | 17.4 |
| BioMysteryBench, biology (hard) | 49.4 | 46.5 |
| Humanity's Last Exam, reasoning without tools | 56.3 | 56.5 |
| DeepSWE v1.1, agentic coding | 68.8 | 69.7 |
| FrontierCode v1.1, agentic coding | 53.4 | 53.5 |
| Legal Agent Benchmark, legal | 11.7 | 13.3 |
Two notes belong with that count before anyone repeats it.
The first is that one of the eight wins is a walkover. Fable 5 posted no score at all on ARC-AGI-3, so Opus 5 wins that row by default. The stricter count is seven clear head to head wins, one uncontested, and four losses.
The second is that the four losses are close. The margins are 0.2 on reasoning without tools, 0.9 on DeepSWE, 0.1 on FrontierCode, and 1.6 on legal. Three of those four are under a single point. Calling them losses is accurate; calling them gaps is not.
Here is the source table itself, with all four model columns. Click through to read it on Anthropic's page.
Two rows in that image do not compare against Fable 5 at all, which is why they sit outside the twelve above. On HealthBench Professional and on the human solved half of BioMysteryBench, the second column is labelled Mythos 5, not Fable 5.
Where Claude Fable 5 Still Wins
Claude Fable 5 keeps four rows, and two of them are coding. If your day is spent driving autonomous coding agents through long tasks, or working through legal research, Fable 5 still edges Opus 5 on the benchmark. Everything else on this table goes to Opus 5.
That is a narrower recommendation than it sounds, because of the margins above. A 0.1 point difference on FrontierCode is not a reason to pay double. The case for Fable 5 on coding rests on it being genuinely ahead at the very top of the range, not on the size of these particular numbers.
The Cost Story: Same Price as Opus 4.8, Half the Cost of Fable 5
The three charts Anthropic published plot score against cost per task at each effort setting, from low through max. Effort is the dial that decides how much thinking and tool use a model spends on a task, and it is why one line on these charts carries five different prices.

Higher on the chart means better answers. Further left means less money. On Frontier-Bench v0.1, the Opus 5 line sits above both Fable 5 and Opus 4.8 across the visible range, which is the clearest of the three.

CursorBench is where the concession lives, and it is worth making. At the far right of that chart, Fable 5's peak score edges past Opus 5's peak. Anthropic says so directly: on CursorBench 3.2 at max effort, Opus 5 "performs within 0.5% of Fable 5's peak score, but at half the cost per task."
So Fable 5 does reach a higher ceiling on that benchmark. It needs more than double the money per task to get there. For the same spend, Opus 5 is the better answer, and that is the claim worth carrying.
These are Anthropic's own runs, so read them as directional. The Frontier-Bench chart footnotes its setup, a mini-SWE-agent harness on a GKE backend averaged over five attempts per task, and notes that Opus 4.8 stood in whenever a safety classifier refused a request for Opus 5 or Fable 5. That footnote covers that one chart. What makes the set credible is what makes the table credible: Anthropic published four rows where its own cheaper model loses.
Where Claude Opus 5 Is Still Behind
Claude Opus 5 is not the top line on every chart, and anyone telling you it swept the launch has not read them closely.

On the Artificial Analysis Coding Agent Index, the GPT-5.6 Sol line sits at or above Opus 5 across most of the price range, and Sol is the cheapest option at the low end of all three charts. Opus 5 beats Fable 5 there. It does not beat everything.
The same thing happens inside the benchmark table, which is why the four column version of it is worth looking at. On DeepSWE v1.1, one of the rows Fable 5 takes from Opus 5, the actual winner is neither of them. GPT-5.6 Sol tops that row at 72.7, ahead of Fable 5's 69.7 and Opus 5's 68.8.
Inside Anthropic's own line up, Mythos 5 leads on health, scoring 66.0 against 59.8 for Opus 5 on HealthBench Professional. Mythos 5 also leads on cybersecurity and on autonomous biology research.
The accurate summary of the launch is that Opus 5 beats Fable 5 on most rows for half the money. It is not that Opus 5 is the best model at everything.
What the Benchmarks Do Not Show: It Checks Its Own Work
The behavioral change is the part no chart on that page captures, and it is the reason I switched my own default over on day one.
Anthropic's announcement describes Opus 5 as much stronger at verifying its work and iterating carefully until it succeeds, and gives three examples. The most legible one involves a real open source package manager bug: Opus 5 found the root cause and caught an edge case that the community patch had missed, while a competing model fixed the visible symptom and reported the job done. In another, given no way to view a drawing, it wrote its own computer vision pipeline to pull the geometry out of raw pixels, and no competing model with the same setup solved it in five attempts.
I run Claude every day across multiple projects, my own and my clients'. Roam Radar and this site are two of the public ones. The rest are internal builds and client facing work, which is the part that matters here: this is not one hobby repo, it is a working week across codebases I did not all write myself. I am on the Max plan, and Opus 5 became my working default the day it launched.
In the first few hours of running it, what stood out had nothing to do with speed or score. Opus 5 flagged problems I had not asked it to look at.
On one project it went back through the history and pointed out that the approach I was about to take had already been tried, many commits earlier, and had not worked. Then it gave me alternatives and made a recommendation rather than waiting to be asked. No model I had used before did that without being told to go looking.
That is one launch day across my own projects and client work. The benchmark rows above are the evidence; this is the thing the rows do not have a column for.
Why I Start Claude Opus 5 on Low Effort
The setting I would change first is effort, and the direction is the opposite of the one most people reach for. I run Opus 5 on low or medium for most of what I do, and high when the task calls for it.
Go back to the two effort charts above and read the leftmost Opus 5 point, which is low effort, against the rightmost Opus 4.8 point, which is the most expensive setting the older model has. On Frontier-Bench v0.1, Opus 5 at its cheapest sits above Opus 4.8 at its most expensive, for roughly a third of the cost per attempt. On CursorBench the same crossover happens by a thinner margin, for under half the cost per task. The floor of the new model is above the ceiling of the model it replaces.
That matches what I have seen running it. On ordinary work I leave it low, and what comes back is what I would have expected from a high setting on the models I was running the week before, Sonnet 5 included. The difference shows up in usage rather than in quality: the same working day burns noticeably fewer tokens, so the limits arrive later. Medium is where a lot of my real work lands, and I go to high for the tasks that earn it, which is rarer than I assumed it would be.
So if you take one change from this post, take that one. Start low and let the task move the dial up, rather than starting high because the work feels important.
My Take: Who Is Actually Paying for This
A better model, at the same price as the one it replaces. That is a good deal, and I think it is worth asking out loud why the deal exists.
My read is that we are being subsidized, and that this is the most interesting thing about the current pricing. Capability went up and the price did not move. Running these models did not get cheaper in the same week, so somebody absorbed the difference, and it was not the person paying monthly.
Price competition is real, though it is worth being precise about what it has actually done. Capable models now come from more places than they used to, and enough of the newer ones compete on cost rather than on topping a leaderboard that price has become a feature instead of a footnote. What that has not done is bring frontier prices down. The top of the range has gone the other way: Fable 5 sits at $10 and $50 per million tokens, double the tier below it. The pressure shows up as models holding their price while getting better, which is exactly what Opus 5 did, rather than as anything getting cheaper.
So the conclusion I draw is to make the most of this while it lasts. Build the habits and the systems now, while frontier capability is priced like this. Pricing set by a race between a handful of companies is not a permanent state of the world, and the people who will feel a change least are the ones who already know exactly what they use these models for.
Should You Use Claude Opus 5?
For most people on a Claude plan the question is already answered, because Opus 5 is the default on Max and the strongest option on Pro. The real decision is when to reach for something else.
| The work | Reach for | Why |
|---|---|---|
| Agentic coding, search, computer use | Claude Opus 5 | Wins those rows outright and costs half of Fable 5 |
| Everyday work you run all day | Claude Opus 5 starting on low effort | Sits above Opus 4.8's most expensive setting on both effort charts, for a fraction of the spend |
| Long autonomous coding runs at max effort | Claude Fable 5 | Holds a slightly higher ceiling on DeepSWE and FrontierCode |
| Legal research | Claude Fable 5 | Takes the legal row by the widest of the four margins |
| Cost sensitive coding volume | GPT-5.6 Sol | Cheapest at the low end of all three charts |
If you want to spend less without dropping to a weaker model, there is a middle path I have written up before. The advisor setup gets you most of a top model's quality for a fraction of the cost by keeping the cheap model on the work and bringing the expensive one in only for the hard calls. And if you are choosing inside the Claude line up rather than against Fable 5, Sonnet vs Opus covers when the cheaper one is the right call.
The Model Matters Less Than the System Around It
A launch like this one is easy to over read. Which model you run does an outsized amount of the work in whether a task lands, and that is exactly why a table like the one above gets the attention it does.
The part fewer people act on is that the model is the smaller half of the job. What separates two people running the identical model is everything around it: the context you hand it, the instructions it reads before it starts, what it remembers between sessions, and whether the work has been broken into pieces a model can actually finish. I have watched a weaker model inside a well built setup get further than a stronger one being used as a chat box, and it is not close.
So I would not spend long agonizing over which row of that table applies to you. If you are on Max or Pro you are already on Opus 5, and the return on your next hour is higher if you spend it on the system you run the model inside rather than on which model it is. That work also survives the next launch, which will be along shortly. If you want to build that setup properly, the AI workshop covers it live, working in Claude Code side by side.
Frequently Asked Questions
Is Claude Opus 5 better than Claude Fable 5?
On Anthropic's own comparison, Claude Opus 5 wins 8 of 12 benchmark rows against Claude Fable 5 and costs half as much per task. Fable 5 keeps four rows, including two coding benchmarks and legal, though three of those four margins are under a single point. For most work Opus 5 is the better call on price alone; for long autonomous coding runs at maximum effort, Fable 5 still holds a slightly higher ceiling.
How much does Claude Opus 5 cost?
At the time of writing, Claude Opus 5 costs $5 per million input tokens and $25 per million output tokens, which is identical to Opus 4.8, the model it replaces. On a subscription the pricing question does not come up at all: Opus 5 is the default model on Claude Max and the strongest model on Claude Pro, at no change to the plan price.
Is Claude Opus 5 available on the Claude Pro plan?
Yes. Claude Opus 5 is the strongest model available on Claude Pro, and it is the default model on Claude Max. Anyone on either plan is using it without changing a setting, which is why the launch matters more to subscription users than the per token pricing suggests.
What is Claude Opus 5 best at?
Claude Opus 5 is strongest at agentic terminal coding, knowledge work, agentic search, computer use, and business workflows, where it beats Claude Fable 5 by the widest margins on Anthropic's table. Its largest win is agentic terminal coding on Frontier-Bench v0.1, at 43.3 against 33.7. The behavioral pattern Anthropic highlights is that it verifies its own work rather than reporting a task complete.
Should you run Claude Opus 5 on low effort?
Start there, yes. On both of Anthropic's agentic coding effort charts, the low effort Opus 5 point sits above the highest effort Opus 4.8 point while costing less per task, so the cheapest setting on the new model beats the most expensive setting on the one it replaces. Day to day I sit on low or medium depending on the task and go to high when one calls for it, and the saving shows up as tokens rather than as weaker answers. Move the dial up for the tasks that need it instead of starting there.
Does Claude Opus 5 beat GPT-5.6 Sol?
Not on coding. On the Artificial Analysis Coding Agent Index, the GPT-5.6 Sol line sits at or above Claude Opus 5 across most of the price range, and Sol is the cheapest option at the low cost end of all three of Anthropic's effort charts. Opus 5's win is against Claude Fable 5, not against every model on the chart.
The Bottom Line on Claude Opus 5
Claude Opus 5 wins 8 of 12 rows against Claude Fable 5, at the same price as Opus 4.8 and half of Fable 5's cost per task. The four losses are real and three of them are inside a point. It is not the top model on every chart, and Anthropic did not claim it was. For anyone on Max or Pro the model behind your work already changed, so the useful question is what to stop sending elsewhere, and which effort setting to start it on. Mine is low, and the task moves it up from there. I will be running it daily and writing up where it holds up and where it gets things wrong.

