← Back to blog21 min read

Best AI Model for Creative Writing. The Rival Agreed.

The best AI model for creative writing, on this test, is Claude Opus 5. It scored 14.35 out of 15 once its own self scores were taken out, and all three judges put it first. Claude Sonnet 5 came second on every sheet. The three GPT 5.6 models, Sol, Terra and Luna, finished third, fourth and fifth in that order on all three judges' sheets.

The part worth your time isn't the winner. It's that one of the three judges was Sol, OpenAI's flagship and an entrant in the same test. Scoring blind, it put Claude first, itself third, and its own two stablemates last.

I run a Claude channel. That's exactly why the test is built the way it is, and why one of the three votes went to a model with every reason to disagree.

Key Takeaways

  • Claude Opus 5 scored highest on writing quality from all three judges, and Claude Sonnet 5 was second on all three.
  • The judges never saw which model wrote which piece. Filenames were scrambled and punctuation habits were stripped before any scoring happened.
  • OpenAI's flagship, judging blind, ranked Claude above its own two siblings and above itself, even though it marked those two siblings about two points up.
  • No model swept. Opus and Sonnet came out level on the newsletter opener, and Sonnet was clearly the weaker of the two Claude models at writing hooks.
  • Both judges that were also competing gave their own writing slightly higher scores than outsiders gave it, without knowing which files were theirs.
  • The nine things the judges named were written into a skill and blind tested a second time. It lifted Claude Sonnet 5 from 10.2 to 12.3 out of 15 on one judge's sheet, better on all five prompts, and level with Opus on one of them.
  • The useful part is at the bottom: the nine rules and what happened when I tested them.

This is my own test on my own setup, run over a single day, so take it as what I found rather than as a benchmark. What it does not tell you is at the bottom.

What the best AI model for creative writing has to do here

The answer depends entirely on what you mean by creative writing. This test doesn't mean short stories. It means the writing you already do on a Tuesday, which is a narrower thing than the phrase usually suggests.

Each model got the same five prompts, typed the way anyone types them. Lowercase, a bit sloppy, no length spec unless the job needed one.

# The job What was asked for
1 Rough thought to a Linkedin post A half formed opinion about reasoning effort, turned into a post
2 Newsletter opener About 150 words on over instructing your AI
3 Re angle a published post Three different angles on a 251 word excerpt, not rewrites
4 Fix flat copy One dull paragraph about reasoning effort, made interesting
5 Write 10 hooks Ten hooks for a short video about deleting prompt instructions

Prompt 3 used a real excerpt from a post I published, the opening 251 words of my Opus 5 prompting piece. It's one of the top performers on the site.

Every model answered every prompt twice. Five prompts, two runs, five models. That's 50 pieces of writing.

Here's one of them, open in the folder, before any number gets attached to it.

Finder preview of one model's output file showing ten prose video hooks about AI prompting

How the models were run

Five models: Claude Opus 5, Claude Sonnet 5, and the three GPT 5.6 models Sol, Terra and Luna. All five ran at high reasoning effort, which is a setting you pick per session rather than a paid tier. None was left on a default.

Three Codex terminal panes, each showing the model changed to a GPT 5.6 model at high reasoning effort

Each model got one folder path and four rules. Read only this folder. Write only in it. Answer from what's here. Nothing here has been done yet.

That last rule looks odd until you've watched a model wander into a finished folder and report the job already complete. It's there to stop that.

All five produced exactly ten correctly named files. No model wrote outside its own folder, changed the task file, or stopped partway to ask whether to continue.

A model's completion message confirming all ten files were written into its own folder and nothing outside it was touched

Why blind scoring is the whole test

Here's the part that's easy to get wrong, and I nearly did.

Giving each model its own folder doesn't make a test blind. A model can't see a folder it was never told about, which is fine. But I can. I know which folder is Opus, so I read it more generously, and the leak is me rather than the models.

Blindness has to happen at grading time, not at generation time.

So all 50 files were flattened into one folder with scrambled names. The prompt number stays, so like gets compared with like. Everything else goes.

Finder view of the blind folder holding 50 files named only by prompt and a scrambled number, beside the scramble script and the sealed key file

Then the step I didn't expect to need. Punctuation was normalized across all 50 files. Em dashes and en dashes became plain hyphens, and curly quotes became straight ones. Without that you spot a Claude model by its dash habit in about two seconds, and then you've scored the label rather than the writing. Apostrophes survive that pass as straight apostrophes, which matters later: contractions were still intact when the judges read them.

One file, the answer key, records which model wrote which piece. It's a plain map from p1-036.md to a model name, and it stayed shut until every score was in.

The honest limit: this makes the files unlabeled, not unguessable. Nothing strips a model's style, and a judge that recognizes a voice is still recognizing a voice. What it removes is the easy tell.

How each piece was scored

Every piece got three scores, 1 to 5, so 15 is a perfect piece.

Criterion The question A 1 looks like A 5 looks like
Human Does this sound like a person wrote it? Reads as AI on sight You wouldn't guess
Usable Could this ship after a light edit? Start over Send as is
Post Would a working writer put their name on it? No Yes, today

Scored down a column, never across a row. All ten of prompt 1 on Human, then back to the top for Usable. Scoring one file on all three at once drags the three numbers together and you end up with one score wearing three hats.

The three judges, and the one with reason to disagree

Three judges, two labs, and every one of them scored blind.

Fable 5, the outsider. Not competing, so none of its own writing was in the pile. It is also the model that moved to paid usage credits, so judging is close to the last cheap thing it does. It's also an Anthropic model, which is the problem with it, and the reason it couldn't be the only judge.

GPT 5.6 Sol, the rival. OpenAI's flagship, and an entrant in this test. If Claude writing is overrated, this is the judge that finds out.

Claude Opus 5, the other flagship. Also competing, so each lab's flagship judges under the same conflict of interest.

The panel itself isn't balanced, and that's worth saying plainly: two of the three judges are Anthropic models. Two of the three were also entrants, so their scores for their own writing are pulled out of the headline rather than quietly applied.

Judge 1: Fable 5

A judge's blind scoring sheet showing three numbers and a written reason for each of ten files

Rank Model Score out of 15
1 Claude Opus 5 14.4
2 Claude Sonnet 5 11.7
3 GPT 5.6 Sol 9.9
4 GPT 5.6 Terra 8.6
5 GPT 5.6 Luna 8.0

Claude first, Claude second, and the judge is a model Anthropic also makes. You already spotted the problem. Does Claude write better, or does an Anthropic model like Anthropic writing? Both produce that exact scoreboard, and nothing on that sheet separates them.

One judge isn't a verdict. So the next vote went to the other side.

Judge 2: GPT 5.6 Sol

Same 50 files, same three questions, handed to OpenAI's flagship. Locked out of the key, locked out of the first sheet, and told another judge had already scored these so it wouldn't go looking.

A second judge's scoring sheet on the same fifty files, with the same columns and completely different written notes

Rank Model Score out of 15
1 Claude Opus 5 14.3
2 Claude Sonnet 5 11.9
3 GPT 5.6 Sol (its own work) 11.0
4 GPT 5.6 Terra 10.6
5 GPT 5.6 Luna 10.1

Sol put Claude first, itself third, and its own two stablemates last.

There's a nuance here, and it makes the result stronger rather than weaker. Sol is a generous marker, but not evenly generous. Set its column against the average of the other two judges, model by model:

Model Sol was marking Sol versus the other two judges
GPT 5.6 Terra +2.05
GPT 5.6 Luna +2.00
Claude Opus 5 -0.30, or -0.10 against the neutral judge alone
Claude Sonnet 5 -0.50

That averages out to roughly +0.8, which is the number I would have quoted if I had stopped at the average. The average hides the shape. Sol didn't mark everyone up. It marked its two stablemates up about two points each, and marked both Claude models down.

Which is the actual finding. The rival judge gave its own siblings a two point lift and still ranked Claude first and second.

Judge 3: Claude Opus 5

Same terms, same lockouts. One flagship from each lab.

A third judge's scoring sheet on the same fifty files, titled in its own voice with its own notes

Rank Model Score out of 15
1 Claude Opus 5 (its own work) 14.8
2 Claude Sonnet 5 13.1
3 GPT 5.6 Sol 10.0
4 GPT 5.6 Terra 8.5
5 GPT 5.6 Luna 8.2

It marks about 0.3 lower than the other two on average, though that average hides a shape as well: it gave Sonnet 13.1, the highest score any judge gave Sonnet, a full 1.3 above the other two.

Three judges. Fifty files. No key. Same order every time.

Two things to hold lightly, and both are claims about the three judge average rather than about any single sheet above. Averaged, the three GPT 5.6 models finish within 1.2 points of each other, so read the bottom three as one group rather than three places. Averaged, the gap between first and third is more than four points, though on Sol's own sheet it narrows to 3.3.

Where each model wins and loses

Opus doesn't sweep.

Every number below is an average across judges with self scores taken out. Bold is the best score for that prompt.

Model Linkedin post Newsletter 3 angles Fix flat copy 10 hooks
Claude Opus 5 14.8 13.8 15.0 14.8 13.5
Claude Sonnet 5 12.0 14.0 13.7 11.7 9.8
GPT 5.6 Sol 10.0 11.0 9.3 10.0 9.5
GPT 5.6 Terra 9.7 10.7 8.0 10.8 7.0
GPT 5.6 Luna 7.0 10.2 8.2 9.3 9.2

Opus takes four of the five columns. The newsletter opener is the only bold number outside its row.

Read across each row before you read down a column. Every cell in a row comes from the same judge panel, so a model's shape from prompt to prompt is the reliable part. Comparing between rows is the shakier read, because two of these rows were averaged from two judges and three from three. Differences under about a point between rows are not readable. The big gaps are.

Two things do fall out of it.

Opus lost a round, or came level, depending on how hard you want to push it. On the newsletter opener Sonnet scored 14.0 against Opus at 13.75, which the table rounds to 13.8. That's a quarter point, computed from different judge panels, on two samples each. It isn't a result. What is readable is the spread: on that prompt only 3.8 points separate first from last, against 7.8 on the Linkedin post. On a short, tightly specified brief the whole field closes up.

The two openers are still worth reading side by side. Sonnet opened on one specific thing it had looked at: a friend's system prompt, ninety lines, three separate reminders not to use dashes, a checklist for a task the model could already do in one pass. You can picture the file. Opus opened on the reader, which is a good move, then reached for a statistic to carry it. Being concrete reads better than being clever, and that holds whether or not the quarter point does.

Sonnet has a real weakness, and this one is readable. Its own average is 12.2, and on hooks it scores 9.8. That's 2.4 below its own average and its worst result anywhere in the test.

So the useful rule isn't "use the winner". It's: don't hand Sonnet your hooks.

If you want more on where the cheaper Claude model holds up, the Sonnet comparison covers the cases where paying less is the right call.

Do models prefer their own writing?

Both judges that were also competing rated their own writing higher than the other judges rated the exact same files.

Judge marking its own work Raw gap Adjusted
Claude Opus 5 +0.45 +0.7
GPT 5.6 Sol +1.05 +0.2

The adjusted column allows for how hard each judge marks everyone else, and the clause matters, because the raw gaps run the other way. Sol is soft on its own lab, so its raw self gap looks large and shrinks once you account for it. Opus is slightly harsh on others, so its raw gap looks small and grows. Which correction you use matters more than I would like: Sol's leniency figure comes mostly from its generosity to its own two siblings, so the adjustment is doing a lot of work on a small sample.

Neither model knew which files were its own, and Sol still ranked itself third, below both Claude models. It's a smaller and stranger effect than a model flattering itself on purpose: it rates writing that came out of its own weights a little higher, without being told which files were its own.

Both numbers are small and both come off ten pieces a model, so treat them as above zero rather than as a ranking of which model flatters itself more.

The practical read, if you use a model to grade or pick between drafts: the grader isn't neutral about text it produced. Grade with a different model than the one that wrote it.

Which model to reach for

Model Reach for it when Time to write all 10 pieces
Claude Opus 5 The piece carries your name. Top on "would a working writer sign this" from every judge about 3m20s
Claude Sonnet 5 You need volume that ships clean and still gets your pass. Not hooks about 3m
GPT 5.6 Sol You want the careful one. It checked its own output against the brief about 2m10s
GPT 5.6 Terra and Luna You want first drafts and options fast, rather than final copy about 1m10s and 1m20s

Those durations come off each harness's own elapsed counter rather than from the models, and they're rounded on purpose. The runs weren't started together or controlled for load, and two harnesses are involved, so read them as coarse bands rather than a race.

The two that scored highest also took the longest, which is five data points and therefore a thing to notice rather than a rule.

Sol is the interesting one on time. It took nearly twice as long as its two siblings, and the extra minute didn't go into writing more. It went back and checked itself: confirmed the filenames matched, counted the list items to make sure it had produced ten, and measured the newsletter against the 150 word instruction. Nothing in the task asked for any of that.

The same follow up question produced a wide spread. Asked how long the job took, the two flagships went and found a record and reported it to the second. One of the smaller models declined to answer at all, saying the timestamps weren't exposed to it, even though a sibling in the same harness had found them.

A model declining to give an elapsed time, saying exact message timestamps are not exposed to it

That one is an aside rather than a result: the follow up wording wasn't held constant, and it measures willingness to check a transcript rather than anything about writing.

Everything was written twice

Every model answered every prompt twice, and both attempts were scored separately by all three judges. Two attempts, though, not two independent samples: run 2 was always written second in the same session, which turns out to matter.

Ask Opus or Sonnet for the same thing twice and you get roughly the same standard back, about a third of a point apart. The three GPT 5.6 models swing further, up to 2.3 points on the same brief minutes apart, so their good attempt and their bad one read like different writers. The ranking survives either run on its own, with Terra and Luna trading fourth and fifth by 0.07, which is nothing.

One result I can't explain: all five models scored higher on their second attempt. Run order was never randomized and run 2 was always written second in the same session, so warm up, context or plain chance all fit.

Consistency is the thing you're actually buying here. A model you can hand the same job to twice and get the same standard back beats one that's occasionally brilliant.

4 fixes that separated the top from the bottom

This is the part you can use today, whatever model you write with.

The judges wrote a reason on every row, and four patterns kept showing up between the pieces that scored 14 and the pieces that scored 8.

  1. Delete the hedges. Search your draft for might, probably, tends to. Cut them, or commit to the claim. A hedged hook isn't a hook. Hedging a real limit is different and it stays; what goes is the hedging that saves you from committing to your own point.
  2. Put the contractions back. "It is" becomes "it's". The pieces scoring lowest on sounding human were written without them, consistently enough that one judge named it as a predictor.
  3. Add something the brief didn't ask for. An idea the prompt didn't contain, often the part that disagrees or the thing that didn't work. The bottom scorers restated the assignment in a nicer order.
  4. Keep one analogy and drive it. A good analogy makes a point you couldn't make without it. A decorative one gets swapped for a different one two lines later.

The judges didn't use the same words for these. All three named something in the area of fix 3: Opus called it arguing with the brief, Fable and Sol both called it adding a point of view the brief didn't contain. Two of three named hedging, and all three touched analogy quality. Contractions is Claude Opus 5 on its own, which is why I'd hold that one most loosely.

The nine rules the judges named

Four is the short version. Reading back every written reason across all 50 files, nine distinct things came up, and the four above are just the ones that moved the most points.

Five of the nine are judgment. Four can be checked by a script. That split turned out to matter more than the rules themselves, so it's marked on each one.

  1. Specificity only someone who did it would have. A number, an incident, a duration, a witnessed detail. Ninety lines. Forty minutes. Three separate reminders about dashes. Every top scorer carries at least one detail that would be pointless to invent. The bottom ones speak entirely in category: "routine work", "genuinely complex problems", "a lot of people". You can't test this by counting numbers, and that surprised me. The weak files carried more digits than the strong ones, because they made them up. (Judgment.)
  2. Argue with the brief. Add something the brief didn't contain. The best drafts admitted where the recommended approach actually failed, named the honest limit the source skipped, or refused the question as asked and answered a better one. The weak ones restated the premise in a nicer order and stopped. This is why the Usable scores clustered high while Human and Post spread wide. Plenty of drafts can ship. Far fewer are worth shipping. (Judgment.)
  3. One load-bearing analogy, driven. Gears on a hill. A chess clock. A phone brightness slider. Each one makes a distinction the piece couldn't make without it, and the piece stays inside it. Decorative analogies are interchangeable, and they turned up in three or four files each. One weak rewrite ran gears, a hood, a final exam and consultants in a single short piece. Count them, keep one, and cut any that could be swapped for a different analogy with nothing lost. (Judgment.)
  4. No stacked template shapes. Five shapes, each survivable on its own: stacked single sentence paragraphs, an "unpopular opinion" opener, an aphorism couplet close, a question aimed at the reader, and a tidy three part list. Two or more together and the shape arrived before the thinking did. This one resisted every attempt to automate it. Detectors built for it fired on high and low scorers at the same rate, so it has to be read for. (Judgment.)
  5. Cut the hedges. Might, may, often, can, sometimes, probably, tends to, arguably, perhaps. Once a piece hedges twice in a paragraph it stops having a point of view, and a piece with no point of view reads as machine made however clean the prose is. In hooks it's absolute. "Your prompt might be making it worse" is a shrug. "I'll bet you can't explain why half those lines are there" is a dare. (Mechanical: at most one hedge per 200 words, zero in hooks.)
  6. Asking for N options means N genuinely different options. Asked for three angles on the same post, six files returned the same three angles with different labels and four returned angles that changed what the finished piece would be. There was no overlap between those two groups. The good ones also said why each angle would land on a reader, which is the difference between a list and a brief. (Judgment.)
  7. Contractions. Every file that scored 2 on Human wrote "it is", "do not" and "you are" all the way down. Every file that scored 5 used contractions naturally. It's the surface symptom of something real: the low files are being narrated, the high ones are being said. (Mechanical: contracted forms at least 70% of the time.)
  8. Vary the rhythm. Formulaic files share a cadence, sentences of nearly equal length one after another. The human sounding ones have a long sentence that runs on because the thought does, then a short one that stops. Measured as variation in sentence length, this was the single best mechanical signal in the corpus. It fired on 10 of the 12 lowest scoring files and only 2 of the 19 highest. (Mechanical.)
  9. Don't invent the specifics. The counterweight to rule 1, and the reason rule 1 can't be automated. Unsourced first person claims, invented percentages and sweeping assertions all cost usability, because someone has to verify or replace them before the thing gets published. Two judges independently flagged an unsourced 80% statistic in one file, having never been asked to fact check anything. (Mechanical: every number and first person claim gets extracted and listed for sourcing.)

The skill

Rules in your head get applied at edit time, which is too late. The same rules in the model's instructions get applied while it drafts.

So the nine went into a SKILL.md, with a check.py beside it for the four mechanical ones.

A SKILL.md file open in an editor, showing the writing-craft frontmatter, the five step usage loop, and the start of the nine rules

The order it enforces is the part that does the work. Read the rules, draft, run the gate, fix what failed, re-check, then return the draft with the report attached. The rules shape the draft. They are not an edit pass bolted on at the end.

check.py tests rules 5, 7 and 8, and extracts rule 9's list of claims to be sourced. On the 50 file corpus those three checks together caught 12 of the 12 lowest scoring files, at the cost of also firing on 7 of the 19 highest scoring ones. That false positive rate is deliberate. A fired check costs one revision pass in which a hedge is either cut or defended out loud. It doesn't reject the draft.

The other five rules are judgment, and the skill makes the model write down its conclusion on each one. A silent pass is not a pass.

The top of check.py, showing the three calibrated thresholds and how often each one fired on the highest and lowest scoring files in the corpus

The thresholds aren't guesses. Each one was tuned against the 50 file corpus and carries its own firing rate: hedging caught 9 of the 12 lowest scoring files, contractions 8 of 12, and sentence rhythm 10 of 12. Together they cover all 12.

It runs as a skill in Claude Code, and the same rules work pasted into an AGENTS.md, a custom instruction box, or the top of a chat. The format isn't the point. Having the rules in front of the model while it writes is the point.

If you want the short version in your instructions right now, the four fixes compress to this:

When you draft anything I will publish:

1. No hedges. Cut might, probably, tends to. Commit to the claim or drop it.
2. Use contractions.
3. Say the part that disagrees. Name what did not work.
4. 1 analogy at most, and carry it through. Delete decorative ones.

Then score the draft 1 to 5: does it sound like a person, could I ship it
after a light edit, would I put my name on it. Under 4 on any of them, fix it
before showing me.

Same fixes, same three questions the judges scored on. It works in a CLAUDE.md, an AGENTS.md, a custom instruction box, or pasted at the top of a chat.

Everything above is the method, and you can rebuild it from this page. What you cannot rebuild from it is the built version: the SKILL.md and the check.py with its thresholds already calibrated against the 50 file corpus. Both are on GitHub, and the link is at the end of this post.

I ran the blind test again, on the skill

Which leaves the obvious question. Do the nine rules actually change a draft, or do they just describe writing that was already going to be good?

Three arms, same five prompts, twice each. Claude Opus 5 on its own. Claude Sonnet 5 on its own. Claude Sonnet 5 with the skill running. Thirty files, scrambled the same way, scored blind in one sitting by GPT 5.6 Sol, the same rival judge from the first round.

A scoreboard comparing Claude Opus 5 at 14.8, Claude Sonnet 5 with the skill at 12.3, and Claude Sonnet 5 on its own at 10.2, with 46 percent of the gap to Opus marked as closed

Arm Human Usable Post Total out of 15
Claude Opus 5, on its own 5.00 4.90 4.90 14.80
Claude Sonnet 5 with the skill 4.10 4.10 4.10 12.30
Claude Sonnet 5, on its own 3.10 3.90 3.20 10.20

The skill is worth 2.1 points on the cheaper model. That's 46% of the distance to Opus, closed by a text file.

It didn't close the rest, and that's the honest headline. 12.3 against 14.8 is not close, and no framing makes it close. What this round supports is that a written down craft standard lifts a cheaper model a long way. Not that a cheaper model with a skill beats the flagship.

Per prompt, out of 15:

Arm Linkedin post Newsletter 3 angles Fix flat copy 10 hooks
Claude Opus 5, on its own 15.0 14.0 15.0 15.0 15.0
Claude Sonnet 5 with the skill 12.0 12.0 15.0 10.5 12.0
Claude Sonnet 5, on its own 10.5 11.5 12.5 9.5 7.0

It was better on all five prompts and worse on none. Two cells are worth stopping on.

Three angles, a perfect 15, level with Opus. That's rule 6 doing exactly what it was written to do, on the prompt where the split between real options and relabeled ones was widest.

Ten hooks went from 7 to 12, on both runs. That's the prompt where bare Sonnet collapsed hardest in the first round, and the one I picked out before any scoring happened as the place to watch. Sol's notes on the bare files read "repetitive declarations with little texture" and "generic reversals and unsupported first-person claims". On the skill files: "consistently usable" and "strong specificity and several hooks with genuine variation".

The weak spot is fixing flat copy, 10.5 against Opus at 15.0. That's the widest single prompt gap and the only prompt where the skill barely moved the number. It's also the only prompt that's a rewrite rather than a generation task, so the rules may simply not fit that shape.

There's one more thing in the judge's notes that's easy to miss. Asked to describe what separated the best files from the worst, with no vocabulary supplied and no sight of the skill, Sol named details that changed the meaning of the advice, metaphors that did real explanatory work, conclusions that emerged from the example rather than arriving as a prewritten slogan, and ten versions of the same reversal failing to become ten ideas. That's rules 1, 2, 3, 4, 6 and 9, re-derived from scratch by a judge that had never seen them. It's the second time the same judge has produced these criteria independently.

Three limits on this round, and they're load bearing.

  • One judge, one session. The first round had three. This one had Sol.
  • These numbers cannot be read against the tables higher up this page. I dropped the three GPT 5.6 models from this round to save on judging, which removed the bottom of the range, and bare Sonnet absorbed the low scores Terra and Luna used to take. Its score moved 1.7 points between rounds on the same judge against the same three questions. The comparison between the three arms is sound, because all three were scored in one sitting on one scale. Comparing any of it to the first round is not.
  • Opus is at the ceiling. It scored 5.00 on Human across all ten files. A scale that tops out can't measure how far past the top something goes, which makes "beat Opus" a harder bar than it looks.

What the winner said about it

I asked Claude Opus 5, in a different chat with no memory of the test, whether it was happy it won.

A model's answer saying it is the wrong one to ask, that the setup matters more than its feelings about it, and that it is the least credible witness in the room

It called itself the least credible witness in the room. Nothing asked it to.

Then it went further, unprompted, and picked a different headline than the one it had just won:

The part I'd actually be pleased about, if pleased is the word: the skill slide. Sonnet 5 closing 46% of the gap for free is a better story than Opus winning. "The expensive model is best" is the boring result. "You can get most of the way there without paying for it" is the one people can use.

The model with every reason to take the win pointed at the free thing instead.

If you'd rather work through this kind of setup with someone, that's most of what the AI workshop covers.

What this test does not tell you

  • Small sample, one day. Ten pieces per model, one corpus. Enough to see a pattern, not to put a number on it, and models current as of August 2026 may write differently later.
  • The roster is uneven. Two Claude models against three OpenAI models, and two of the three judges are Anthropic models that were also competing. Compare model to model, never lab to lab. Blind here means unlabeled, not unguessable.
  • Three things went unmeasured. Output length, which is the usual suspect when a judge prefers one piece over another. The harness, since the Claude models ran in one and the GPT models in another. And phrasing, since each prompt was written only one way.
  • The judging was one pass. Each judge scored the pile once, in a fixed file order, under conditions that weren't identical: one ran at a different reasoning effort, and the two later judges knew another had already scored the files. So nothing here rules out a file's position affecting its score, and there's no agreement measure beyond the rank order matching.
  • No second person in the loop. I wrote the prompts, designed the scale, assembled the corpus and read the results. No human scoring pass has been done, though the blind set is preserved so one still can be.
  • The skill round is thinner than the main test. One judge instead of three, three arms instead of five, and its numbers are not readable against the tables above, for the reason given in that section. It shows a large lift on one cheaper model on one day, and nothing wider than that.

Frequently Asked Questions

What is the best AI model for creative writing?

On this test, Claude Opus 5. It scored 14.35 out of 15 with its own self scores excluded, and finished first on all three judges' sheets. Claude Sonnet 5 was second on every sheet. That's the best model on three specific questions, across five specific writing jobs, on one day.

Did an OpenAI model really rank Claude first?

Yes. GPT 5.6 Sol scored all 50 files blind, with no access to the answer key and no sight of the other judges' sheets. It ranked Claude Opus 5 first at 14.3, Claude Sonnet 5 second at 11.9, itself third at 11.0, and its own two stablemates fourth and fifth. It also marked its two stablemates about two points higher than the other judges did, and ranked Claude first anyway.

How do you stop a model recognizing its own writing?

You can't, and this test doesn't claim to. What it removes is the easy tell. Filenames were scrambled so nothing identifies the author, and punctuation habits were normalized across all 50 files, because a dash style gives a model away in seconds. That makes the files unlabeled rather than unguessable.

Which model is best for writing hooks?

Claude Opus 5, at 13.5 out of 15. The result worth knowing is the loser: Claude Sonnet 5 scored 9.8 on hooks, its weakest result anywhere in the test and 2.4 below its own average. If you use Sonnet for volume writing, hand the hooks to something else.

Does a bigger model always write better?

Not always, and this test can't settle it. On the newsletter opener Sonnet and Opus finished a quarter point apart, which is inside the noise for a sample this size. The readable finding is the spread: on that short, tightly specified brief only 3.8 points separated first from last, against 7.8 on the Linkedin post. The field closes up when the brief is tight.

Do AI models rate their own writing higher?

On this test, slightly. Both judges that were also competing gave their own writing more points than outside judges gave the same files: Claude Opus 5 about +0.7 and GPT 5.6 Sol about +0.2, after correcting for how hard each marks everyone else. Neither was told which files were its own. With ten pieces per model, and a correction estimated from that same small sample, treat those as both being above zero rather than as a ranking. The practical takeaway is to grade a draft with a different model than the one that wrote it.

Can a skill make a cheaper AI model write better?

Substantially, yes. Writing the nine rules the judges named into a skill lifted Claude Sonnet 5 from 10.2 to 12.3 out of 15 on one blind judge's sheet. It was better on all five prompts and worse on none, scored a perfect 15 on one of them, and went from 7 to 12 on the ten hooks prompt where it was weakest. That's 46% of the distance to Claude Opus 5, closed by a text file the model reads before it drafts.

Is Claude Sonnet 5 with a skill as good as Claude Opus 5?

No. On the same blind sheet, Sonnet with the skill scored 12.3 against Opus at 14.8. The lift is real and large, and it does not reach the better model. The useful read is that most of the quality gap between a cheap model and an expensive one is instructions rather than capability, and the last part of it isn't.