← Back to blog13 min read

Best AI Model for Creative Writing. The Rival Agreed.

The best AI model for creative writing, on this test, is Claude Opus 5. It scored 14.35 out of 15 once its own self scores were taken out, and all three judges put it first. Claude Sonnet 5 came second on every sheet. The three GPT 5.6 models, Sol, Terra and Luna, finished third, fourth and fifth in that order on all three judges' sheets.

The part worth your time isn't the winner. It's that one of the three judges was Sol, OpenAI's flagship and an entrant in the same test. Scoring blind, it put Claude first, itself third, and its own two stablemates last.

I run a Claude channel. That's exactly why the test is built the way it is, and why one of the three votes went to a model with every reason to disagree.

Key Takeaways

  • Claude Opus 5 scored highest on writing quality from all three judges, and Claude Sonnet 5 was second on all three.
  • The judges never saw which model wrote which piece. Filenames were scrambled and punctuation habits were stripped before any scoring happened.
  • OpenAI's flagship, judging blind, ranked Claude above its own two siblings and above itself, even though it marked those two siblings about two points up.
  • No model swept. Opus and Sonnet came out level on the newsletter opener, and Sonnet was clearly the weaker of the two Claude models at writing hooks.
  • Both judges that were also competing gave their own writing slightly higher scores than outsiders gave it, without knowing which files were theirs.
  • The useful part is at the bottom: 4 fixes that separated the top from the bottom.

This is my own test on my own setup, run over a single day, so take it as what I found rather than as a benchmark. What it does not tell you is at the bottom.

What the best AI model for creative writing has to do here

The answer depends entirely on what you mean by creative writing. This test doesn't mean short stories. It means the writing you already do on a Tuesday, which is a narrower thing than the phrase usually suggests.

Each model got the same five prompts, typed the way anyone types them. Lowercase, a bit sloppy, no length spec unless the job needed one.

# The job What was asked for
1 Rough thought to a Linkedin post A half formed opinion about reasoning effort, turned into a post
2 Newsletter opener About 150 words on over instructing your AI
3 Re angle a published post Three different angles on a 251 word excerpt, not rewrites
4 Fix flat copy One dull paragraph about reasoning effort, made interesting
5 Write 10 hooks Ten hooks for a short video about deleting prompt instructions

Prompt 3 used a real excerpt from a post I published, the opening 251 words of my Opus 5 prompting piece. It's one of the top performers on the site.

Every model answered every prompt twice. Five prompts, two runs, five models. That's 50 pieces of writing.

Here's one of them, open in the folder, before any number gets attached to it.

Finder preview of one model's output file showing ten prose video hooks about AI prompting

How the models were run

Five models: Claude Opus 5, Claude Sonnet 5, and the three GPT 5.6 models Sol, Terra and Luna. All five ran at high reasoning effort, which is a setting you pick per session rather than a paid tier. None was left on a default.

Three Codex terminal panes, each showing the model changed to a GPT 5.6 model at high reasoning effort

Each model got one folder path and four rules. Read only this folder. Write only in it. Answer from what's here. Nothing here has been done yet.

That last rule looks odd until you've watched a model wander into a finished folder and report the job already complete. It's there to stop that.

All five produced exactly ten correctly named files. No model wrote outside its own folder, changed the task file, or stopped partway to ask whether to continue.

A model's completion message confirming all ten files were written into its own folder and nothing outside it was touched

Why blind scoring is the whole test

Here's the part that's easy to get wrong, and I nearly did.

Giving each model its own folder doesn't make a test blind. A model can't see a folder it was never told about, which is fine. But I can. I know which folder is Opus, so I read it more generously, and the leak is me rather than the models.

Blindness has to happen at grading time, not at generation time.

So all 50 files were flattened into one folder with scrambled names. The prompt number stays, so like gets compared with like. Everything else goes.

Finder view of the blind folder holding 50 files named only by prompt and a scrambled number, beside the scramble script and the sealed key file

Then the step I didn't expect to need. Punctuation was normalized across all 50 files. Em dashes and en dashes became plain hyphens, and curly quotes became straight ones. Without that you spot a Claude model by its dash habit in about two seconds, and then you've scored the label rather than the writing. Apostrophes survive that pass as straight apostrophes, which matters later: contractions were still intact when the judges read them.

One file, the answer key, records which model wrote which piece. It's a plain map from p1-036.md to a model name, and it stayed shut until every score was in.

The honest limit: this makes the files unlabeled, not unguessable. Nothing strips a model's style, and a judge that recognizes a voice is still recognizing a voice. What it removes is the easy tell.

How each piece was scored

Every piece got three scores, 1 to 5, so 15 is a perfect piece.

Criterion The question A 1 looks like A 5 looks like
Human Does this sound like a person wrote it? Reads as AI on sight You wouldn't guess
Usable Could this ship after a light edit? Start over Send as is
Post Would a working writer put their name on it? No Yes, today

Scored down a column, never across a row. All ten of prompt 1 on Human, then back to the top for Usable. Scoring one file on all three at once drags the three numbers together and you end up with one score wearing three hats.

The three judges, and the one with reason to disagree

Three judges, two labs, and every one of them scored blind.

Fable 5, the outsider. Not competing, so none of its own writing was in the pile. It's also an Anthropic model, which is the problem with it, and the reason it couldn't be the only judge.

GPT 5.6 Sol, the rival. OpenAI's flagship, and an entrant in this test. If Claude writing is overrated, this is the judge that finds out.

Claude Opus 5, the other flagship. Also competing, so each lab's flagship judges under the same conflict of interest.

The panel itself isn't balanced, and that's worth saying plainly: two of the three judges are Anthropic models. Two of the three were also entrants, so their scores for their own writing are pulled out of the headline rather than quietly applied.

Judge 1: Fable 5

A judge's blind scoring sheet showing three numbers and a written reason for each of ten files

Rank Model Score out of 15
1 Claude Opus 5 14.4
2 Claude Sonnet 5 11.7
3 GPT 5.6 Sol 9.9
4 GPT 5.6 Terra 8.6
5 GPT 5.6 Luna 8.0

Claude first, Claude second, and the judge is a model Anthropic also makes. You already spotted the problem. Does Claude write better, or does an Anthropic model like Anthropic writing? Both produce that exact scoreboard, and nothing on that sheet separates them.

One judge isn't a verdict. So the next vote went to the other side.

Judge 2: GPT 5.6 Sol

Same 50 files, same three questions, handed to OpenAI's flagship. Locked out of the key, locked out of the first sheet, and told another judge had already scored these so it wouldn't go looking.

A second judge's scoring sheet on the same fifty files, with the same columns and completely different written notes

Rank Model Score out of 15
1 Claude Opus 5 14.3
2 Claude Sonnet 5 11.9
3 GPT 5.6 Sol (its own work) 11.0
4 GPT 5.6 Terra 10.6
5 GPT 5.6 Luna 10.1

Sol put Claude first, itself third, and its own two stablemates last.

There's a nuance here, and it makes the result stronger rather than weaker. Sol is a generous marker, but not evenly generous. Set its column against the average of the other two judges, model by model:

Model Sol was marking Sol versus the other two judges
GPT 5.6 Terra +2.05
GPT 5.6 Luna +2.00
Claude Opus 5 -0.30, or -0.10 against the neutral judge alone
Claude Sonnet 5 -0.50

That averages out to roughly +0.8, which is the number I would have quoted if I had stopped at the average. The average hides the shape. Sol didn't mark everyone up. It marked its two stablemates up about two points each, and marked both Claude models down.

Which is the actual finding. The rival judge gave its own siblings a two point lift and still ranked Claude first and second.

Judge 3: Claude Opus 5

Same terms, same lockouts. One flagship from each lab.

A third judge's scoring sheet on the same fifty files, titled in its own voice with its own notes

Rank Model Score out of 15
1 Claude Opus 5 (its own work) 14.8
2 Claude Sonnet 5 13.1
3 GPT 5.6 Sol 10.0
4 GPT 5.6 Terra 8.5
5 GPT 5.6 Luna 8.2

It marks about 0.3 lower than the other two on average, though that average hides a shape as well: it gave Sonnet 13.1, the highest score any judge gave Sonnet, a full 1.3 above the other two.

Three judges. Fifty files. No key. Same order every time.

Two things to hold lightly, and both are claims about the three judge average rather than about any single sheet above. Averaged, the three GPT 5.6 models finish within 1.2 points of each other, so read the bottom three as one group rather than three places. Averaged, the gap between first and third is more than four points, though on Sol's own sheet it narrows to 3.3.

Where each model wins and loses

Opus doesn't sweep.

Every number below is an average across judges with self scores taken out. Bold is the best score for that prompt.

Model Linkedin post Newsletter 3 angles Fix flat copy 10 hooks
Claude Opus 5 14.8 13.8 15.0 14.8 13.5
Claude Sonnet 5 12.0 14.0 13.7 11.7 9.8
GPT 5.6 Sol 10.0 11.0 9.3 10.0 9.5
GPT 5.6 Terra 9.7 10.7 8.0 10.8 7.0
GPT 5.6 Luna 7.0 10.2 8.2 9.3 9.2

Opus takes four of the five columns. The newsletter opener is the only bold number outside its row.

Read across each row before you read down a column. Every cell in a row comes from the same judge panel, so a model's shape from prompt to prompt is the reliable part. Comparing between rows is the shakier read, because two of these rows were averaged from two judges and three from three. Differences under about a point between rows are not readable. The big gaps are.

Two things do fall out of it.

Opus lost a round, or came level, depending on how hard you want to push it. On the newsletter opener Sonnet scored 14.0 against Opus at 13.75, which the table rounds to 13.8. That's a quarter point, computed from different judge panels, on two samples each. It isn't a result. What is readable is the spread: on that prompt only 3.8 points separate first from last, against 7.8 on the Linkedin post. On a short, tightly specified brief the whole field closes up.

The two openers are still worth reading side by side. Sonnet opened on one specific thing it had looked at: a friend's system prompt, ninety lines, three separate reminders not to use dashes, a checklist for a task the model could already do in one pass. You can picture the file. Opus opened on the reader, which is a good move, then reached for a statistic to carry it. Being concrete reads better than being clever, and that holds whether or not the quarter point does.

Sonnet has a real weakness, and this one is readable. Its own average is 12.2, and on hooks it scores 9.8. That's 2.4 below its own average and its worst result anywhere in the test.

So the useful rule isn't "use the winner". It's: don't hand Sonnet your hooks.

If you want more on where the cheaper Claude model holds up, the Sonnet comparison covers the cases where paying less is the right call.

Do models prefer their own writing?

Both judges that were also competing rated their own writing higher than the other judges rated the exact same files.

Judge marking its own work Raw gap Adjusted
Claude Opus 5 +0.45 +0.7
GPT 5.6 Sol +1.05 +0.2

The adjusted column allows for how hard each judge marks everyone else, and the clause matters, because the raw gaps run the other way. Sol is soft on its own lab, so its raw self gap looks large and shrinks once you account for it. Opus is slightly harsh on others, so its raw gap looks small and grows. Which correction you use matters more than I would like: Sol's leniency figure comes mostly from its generosity to its own two siblings, so the adjustment is doing a lot of work on a small sample.

Neither model knew which files were its own, and Sol still ranked itself third, below both Claude models. It's a smaller and stranger effect than a model flattering itself on purpose: it rates writing that came out of its own weights a little higher, without being told which files were its own.

Both numbers are small and both come off ten pieces a model, so treat them as above zero rather than as a ranking of which model flatters itself more.

The practical read, if you use a model to grade or pick between drafts: the grader isn't neutral about text it produced. Grade with a different model than the one that wrote it.

Which model to reach for

Model Reach for it when Time to write all 10 pieces
Claude Opus 5 The piece carries your name. Top on "would a working writer sign this" from every judge about 3m20s
Claude Sonnet 5 You need volume that ships clean and still gets your pass. Not hooks about 3m
GPT 5.6 Sol You want the careful one. It checked its own output against the brief about 2m10s
GPT 5.6 Terra and Luna You want first drafts and options fast, rather than final copy about 1m10s and 1m20s

Those durations come off each harness's own elapsed counter rather than from the models, and they're rounded on purpose. The runs weren't started together or controlled for load, and two harnesses are involved, so read them as coarse bands rather than a race.

The two that scored highest also took the longest, which is five data points and therefore a thing to notice rather than a rule.

Sol is the interesting one on time. It took nearly twice as long as its two siblings, and the extra minute didn't go into writing more. It went back and checked itself: confirmed the filenames matched, counted the list items to make sure it had produced ten, and measured the newsletter against the 150 word instruction. Nothing in the task asked for any of that.

The same follow up question produced a wide spread. Asked how long the job took, the two flagships went and found a record and reported it to the second. One of the smaller models declined to answer at all, saying the timestamps weren't exposed to it, even though a sibling in the same harness had found them.

A model declining to give an elapsed time, saying exact message timestamps are not exposed to it

That one is an aside rather than a result: the follow up wording wasn't held constant, and it measures willingness to check a transcript rather than anything about writing.

Everything was written twice

Every model answered every prompt twice, and both attempts were scored separately by all three judges. Two attempts, though, not two independent samples: run 2 was always written second in the same session, which turns out to matter.

Ask Opus or Sonnet for the same thing twice and you get roughly the same standard back, about a third of a point apart. The three GPT 5.6 models swing further, up to 2.3 points on the same brief minutes apart, so their good attempt and their bad one read like different writers. The ranking survives either run on its own, with Terra and Luna trading fourth and fifth by 0.07, which is nothing.

One result I can't explain: all five models scored higher on their second attempt. Run order was never randomized and run 2 was always written second in the same session, so warm up, context or plain chance all fit.

Consistency is the thing you're actually buying here. A model you can hand the same job to twice and get the same standard back beats one that's occasionally brilliant.

4 fixes that separated the top from the bottom

This is the part you can use today, whatever model you write with.

The judges wrote a reason on every row, and four patterns kept showing up between the pieces that scored 14 and the pieces that scored 8.

  1. Delete the hedges. Search your draft for might, probably, tends to. Cut them, or commit to the claim. A hedged hook isn't a hook. Hedging a real limit is different and it stays; what goes is the hedging that saves you from committing to your own point.
  2. Put the contractions back. "It is" becomes "it's". The pieces scoring lowest on sounding human were written without them, consistently enough that one judge named it as a predictor.
  3. Add something the brief didn't ask for. An idea the prompt didn't contain, often the part that disagrees or the thing that didn't work. The bottom scorers restated the assignment in a nicer order.
  4. Keep one analogy and drive it. A good analogy makes a point you couldn't make without it. A decorative one gets swapped for a different one two lines later.

The judges didn't use the same words for these. All three named something in the area of fix 3: Opus called it arguing with the brief, Fable and Sol both called it adding a point of view the brief didn't contain. Two of three named hedging, and all three touched analogy quality. Contractions is Claude Opus 5 on its own, which is why I'd hold that one most loosely.

Hand the fixes to the model instead

The four fixes work by hand. They work better sitting in the model's instructions, so it applies them before you ever see the draft.

## Writing

When you draft anything I will publish:

1. No hedges. Cut might, probably, tends to. Commit to the claim or drop it.
2. Use contractions.
3. Say the part that disagrees. Name what did not work.
4. 1 analogy at most, and carry it through. Delete decorative ones.

Then score the draft 1 to 5: does it sound like a person, could I ship it
after a light edit, would I put my name on it. Under 4 on any of them, fix it
before showing me.

Same four fixes, same three questions the judges used. It works in a CLAUDE.md, an AGENTS.md, a custom instruction box, or pasted at the top of a chat.

If you'd rather work through this kind of setup with someone, that's most of what the AI workshop covers.

What this test does not tell you

  • Small sample, one day. Ten pieces per model, one corpus. Enough to see a pattern, not to put a number on it, and models current as of August 2026 may write differently later.
  • The roster is uneven. Two Claude models against three OpenAI models, and two of the three judges are Anthropic models that were also competing. Compare model to model, never lab to lab. Blind here means unlabeled, not unguessable.
  • Three things went unmeasured. Output length, which is the usual suspect when a judge prefers one piece over another. The harness, since the Claude models ran in one and the GPT models in another. And phrasing, since each prompt was written only one way.
  • The judging was one pass. Each judge scored the pile once, in a fixed file order, under conditions that weren't identical: one ran at a different reasoning effort, and the two later judges knew another had already scored the files. So nothing here rules out a file's position affecting its score, and there's no agreement measure beyond the rank order matching.
  • No second person in the loop. I wrote the prompts, designed the scale, assembled the corpus and read the results. No human scoring pass has been done, though the blind set is preserved so one still can be.

Frequently Asked Questions

What is the best AI model for creative writing?

On this test, Claude Opus 5. It scored 14.35 out of 15 with its own self scores excluded, and finished first on all three judges' sheets. Claude Sonnet 5 was second on every sheet. That's the best model on three specific questions, across five specific writing jobs, on one day.

Did an OpenAI model really rank Claude first?

Yes. GPT 5.6 Sol scored all 50 files blind, with no access to the answer key and no sight of the other judges' sheets. It ranked Claude Opus 5 first at 14.3, Claude Sonnet 5 second at 11.9, itself third at 11.0, and its own two stablemates fourth and fifth. It also marked its two stablemates about two points higher than the other judges did, and ranked Claude first anyway.

How do you stop a model recognizing its own writing?

You can't, and this test doesn't claim to. What it removes is the easy tell. Filenames were scrambled so nothing identifies the author, and punctuation habits were normalized across all 50 files, because a dash style gives a model away in seconds. That makes the files unlabeled rather than unguessable.

Which model is best for writing hooks?

Claude Opus 5, at 13.5 out of 15. The result worth knowing is the loser: Claude Sonnet 5 scored 9.8 on hooks, its weakest result anywhere in the test and 2.4 below its own average. If you use Sonnet for volume writing, hand the hooks to something else.

Does a bigger model always write better?

Not always, and this test can't settle it. On the newsletter opener Sonnet and Opus finished a quarter point apart, which is inside the noise for a sample this size. The readable finding is the spread: on that short, tightly specified brief only 3.8 points separated first from last, against 7.8 on the Linkedin post. The field closes up when the brief is tight.

Do AI models rate their own writing higher?

On this test, slightly. Both judges that were also competing gave their own writing more points than outside judges gave the same files: Claude Opus 5 about +0.7 and GPT 5.6 Sol about +0.2, after correcting for how hard each marks everyone else. Neither was told which files were its own. With ten pieces per model, and a correction estimated from that same small sample, treat those as both being above zero rather than as a ranking. The practical takeaway is to grade a draft with a different model than the one that wrote it.