ICE Scoring: How to Make an AI Agent Rank Your Ideas
ICE scoring ranks a list of ideas on three numbers: Impact, Confidence and Ease. You score each one from 1 to 10, multiply the three together, then divide by 10 to get the ICE score. The highest number is what you build first. It takes about a minute per idea by hand, and your AI agent can score a whole list in minutes if you give it the right instructions. This post covers what each number means, how to score them without guessing, and the five instructions that keep an agent honest while it scores. Jump straight to how to make an agent run it if you already know the framework.
Key Takeaways
- ICE scoring is Impact times Confidence times Ease, each scored 1 to 10, then divided by 10 to give one number per idea.
- Ease is the reverse of effort. A 10 means the work is quick, not that the work is large.
- Tell the agent what you are trying to move before it scores anything, or it scores against its own guess.
- Make it write one line of evidence under every Confidence score. Without evidence, a high Confidence number is just enthusiasm.
- Scoring on Impact alone is the failure that hides in plain sight. Across 9 of my own ideas, my best and my worst both scored 9 on Impact, and one of them came dead last.
What Is ICE Scoring?
ICE scoring is a prioritization method that turns each idea on your list into a single comparable number. ICE stands for Impact, Confidence and Ease, and the result is one number per idea that you sort the list by.
ICE scoring came out of the growth hacking movement. Sean Ellis, the person who coined the term growth hacking, is credited with creating it while leading growth at Dropbox. It stuck because it is fast: three numbers, no reach estimates, no effort in person months.
Why I Keep Coming Back to It
I ran ICE regularly as a growth product manager at a tech startup, which is the environment it was built for: more ideas than anyone could get to, and no shared way to settle which ones went first.
I keep using it because it puts a number on a gut feel without pretending the gut feel is data. Every idea starts as a hunch about how much it matters, and Impact is where that hunch gets written down where it can be compared. Confidence is where it gets checked against what you actually have, and scoring your own certainty is uncomfortable in a useful way, because it makes you say out loud what you actually know. The low Confidence numbers are your blind spots, sitting in a column. Ease then prices the whole thing in time, which is the part a gut feel is worst at.
That combination is what takes the emotion out. The idea I am excited about and the idea I have been avoiding get the same three questions, and once the numbers are on the page the argument is about evidence rather than about who feels more strongly.
What Does Each Letter in ICE Stand For?
| Letter | The question it asks | What a 10 means |
|---|---|---|
| Impact | How much does this move the thing I care about? | It makes a very big difference |
| Confidence | How sure am I that it will actually work? | I have direct evidence I can point to |
| Ease | How easy is this to build? | Under an hour, nothing risky |
A 1 on any of the three is the opposite end. A 1 on Impact means the change barely registers. A 1 on Confidence means you are guessing. A 1 on Ease means the work is large, slow, or likely to go wrong.
Why Divide by 10?
Three numbers from 1 to 10 multiply out to a range of 1 to 1000, which is more digits than anyone needs. Dividing by 10 caps the top of the scale at 100. The ranking is identical either way, so pick a scale and keep it consistent across the whole list.
Ease Is the Reverse of Effort
This is the part of ICE scoring I got backwards the first time I used it. Every other prioritization method scores effort, where a big number means a lot of work. Ease runs the other way. A 10 on Ease means the work is small and quick. A 1 means it is large and slow.
Get it backwards and your ranking inverts: the hardest things float to the top and the fast wins sink. Higher Ease, less work.
How to Score Each Number From 1 to 10
A score is only useful if you would give the same idea the same number tomorrow, which means fixing what 10, 5 and 1 mean before you start.
Impact: How Much Does This Move the Goal?
Impact is scored against one specific thing you are trying to move, not against a general sense of importance. Write that thing down first. A 10 is a very big move on it. A 5 is noticeable. A 1 barely registers.
Impact only means anything against a named goal. An idea that would double your signup rate scores a 10 if signups are the goal and a 3 if you are trying to cut support load this month, and the score is right both times.
Confidence: How Sure Are You That It Will Work?
Confidence is about evidence, not enthusiasm. A 10 means you have direct proof: you have done this before, or you have data showing the effect. A 5 is a reasonable bet. A 1 is a guess dressed up as a plan.
The rule that makes this number honest is simple. Write one line of evidence next to every Confidence score. Where there is nothing to write, the score is a 4 or below. Confidence is the number that quietly drifts upward, because you are scoring your own ideas and you would not have written them down if you did not believe in them.
Ease: How Fast Can You Actually Build It?
Estimate the build time in hours or days first, then convert it to a score. Doing it the other way around produces a number that reflects how you feel about the work rather than how long it takes.
A useful anchor set: under an hour with nothing risky is a 10, a day or two is a 5, and anything large or likely to break something is a 1 or 2.
When Ease comes back low, split the idea. A 2 on Ease usually means the idea is really three ideas wearing one name. Break it into the smallest piece that stands on its own, score that separately, and the piece that was going to sit at the bottom of your list for a year often turns into something you can do this week.
How to Make an AI Agent Run ICE Scoring on Your Backlog
An agent will score a list of twenty ideas in one sitting, and it holds the anchors steadier across a long list than I do by hand, which is the part that slips by the time you reach idea fifteen. What it will not do on its own is score them well. In my own runs, unprompted, it scored against a goal it had guessed at, put Confidence higher than the evidence supported, and handed back a tidy table with nothing underneath it.
Five instructions fix that. They work the same way in Claude, in Codex, or in any agent that can hold a list.
1. Name What You Are Optimizing For, First
Before a single score, tell it the goal metric and the time frame. "Cut the time it takes me to publish a video, over the next 30 days" is a goal. "Improve my pipeline" is not. Handed the second one, mine picked a goal of its own and did not say so.
This is the same discipline as writing a proper prompt brief rather than a request. The scoring is only as good as the target you point it at.
2. Give It the Raw List, Not Your Shortlist
Hand it everything, including the ideas you are fairly sure are bad. Filtering first defeats the point: you are running this because your sense of which ideas are good is what you do not trust.
3. Make It Show Evidence Under Every Confidence Score
Require one line of evidence per Confidence score, and require the words "no evidence" where there is none, with the score capped at 4. This is the instruction I would keep if I could only keep one, because it converts an unfalsifiable number into a claim you can read and disagree with.
4. Make It Estimate Build Time Before It Scores Ease
Ask for the estimate in hours or days as its own column, then let it convert. When I let it score Ease directly, the numbers tracked how complicated the idea sounded rather than how long it would take. An agent that has to commit to "roughly two days" first has to think about the actual work.
5. Have It Argue Against Its Own Ranking
Once the table exists, run a second pass: for the top three, name the strongest reason the Impact score is too high and what has been assumed but not verified. For the bottom three, name what would have to be true for them to jump to the top.
Ask it, in the same pass, to flag any idea scoring 8 or higher on Impact that still lands in the bottom half, and to say what would move it up. That flag is the whole reason the next section exists.
This is the same move as the advisor pattern, where a stronger model reviews a cheaper one's work. A ranking that survives its own author's best attack is worth acting on. One that has never been attacked is just a list you feel good about.

What I Found Scoring 9 of My Own Ideas
I run a content pipeline that turns recordings into finished videos, and I had asked my agent where it could be improved. It came back with 9 ideas. Rather than pick by feel, I had it score all 9 with ICE and hand me back a ranked table.
The top idea scored 72, on Impact 9, Confidence 10 and Ease 8. The bottom came out at 10.8, on Impact 9, Confidence 6 and Ease 2, which my table rounded to 11. The whole list ran from 72 down to 11 with no ties.
My Best and My Worst Idea Both Scored 9 on Impact
That is the finding, and it is the reason I now score all three numbers instead of trusting the one that feels most important.
Impact is the easiest of the three to score and the most satisfying to argue about, so it is the number I reach for when I am in a hurry. On my list, ranking on Impact alone would have put a rewrite I had been talking myself into for weeks near the top. It scored a 9 on Impact and it deserved it. It also scored a 6 on Confidence, against the 10 the top item earned, and a 2 on Ease, because it was a large, slow job with a lot of ways to go wrong. Multiplied out, it came dead last.
The top item was the opposite shape: a small reliability fix on a step that shows up in every video, where I already knew how to do the work and most of the code existed. Unglamorous, and the highest score on the board.
The Ranking Is a Guide, Not a Promise
A low score is a snapshot of what you know today, not a verdict. The bottom item on my list moves up the moment a cheap test raises its Confidence, or I find a way to do a tenth of it and raise its Ease. That is what the second pass is for: it tells you what would have to change, so a low score becomes a to do rather than a rejection.
Two more things ICE scoring will not do for you. It has no term for how many people an idea touches, so a fix that helps everyone and a fix that helps one person can land on the same score. And the scores are yours, which means anyone motivated to see an idea win can score it into first place without lying about a single number.
The instructions above are the method. If you would rather not assemble them yourself, the prompt below is the whole thing written out, both passes and every anchor, ready to paste into your agent as one block.
ICE Scoring vs RICE and the Other Prioritization Methods
RICE is the framework you will run into most often next to ICE, and the difference is what it asks you to know before you can start.
| ICE | RICE | |
|---|---|---|
| Formula | Impact x Confidence x Ease, divided by 10 | Reach x Impact x Confidence, divided by Effort |
| What you need | Three scores from 1 to 10 | Real reach numbers and effort in person months |
| Best for | A backlog you want ranked today | A roadmap with traffic data behind it |
ICE scoring is the fast one, which is why it suits an agent working through a list. RICE is the better answer once you have reliable numbers for how many people an idea touches. Most backlogs never reach that point, and waiting until they do is how a list sits unranked for a year.
What About the Eisenhower Matrix, MoSCoW and Value vs Effort?
Each of those is built for a different job. The Eisenhower Matrix sorts by urgent against important, which is the right tool for deciding what to do today. MoSCoW sorts work into must have, should have, could have and will not have this time, which is how you scope one release rather than how you order the items inside a bucket. Value versus Effort is the closest relative to ICE, and the difference is Confidence: ICE also asks how sure you are of the other two.
For ordering a backlog, ICE is the one I reach for, and Confidence is the reason why.
If you are still deciding what your agent should work on at all, what to decide first covers the ground in front of this one, and I work through prioritization live on real backlogs in the AI Workflow Workshop.
