ICE Scoring — Impact times Confidence times Ease, scored one to ten each. Fast and rough, and blind to how many people an idea reaches.

ICE Scoring: Impact, Confidence and Ease, and When to Use RICE Instead

Sean Ellis c. 2010 Low Complexity

ICE Scoring is a prioritization method that rates each idea on three factors, Impact, Confidence and Ease, and multiplies them into a single score used to rank a list.

Before you start

Is this your framework?

ICE answers one question: given a long list of things we could do, which order should we try them in. It is fast, rough, and honest about being rough.

It is a triage tool. Nothing in it accounts for how many people an idea reaches, for dependencies, or for whether the list contains the right ideas at all. The table below routes to what does.

Matching your actual problem to the right framework.
If your real problem is…You probably want
We need a ranking we can defend to stakeholders, and reach differs a lot between itemsRICE — adds Reach and divides by Effort, slower and more defensible
Compare ICE and RICE
We need to fix what is in and out of a release, not rank a listMoSCoW — categorical scope control for negotiating a delivery boundary
Compare ICE and MoSCoW
We want to know which features customers would actually be delighted byKano Model — researched with customers rather than estimated in a room
We are weighing benefit against work and want it visual, not numericValue vs Effort Matrix — the same trade-off as a two-by-two
The question is where our attention goes this week, not what to buildEisenhower Matrix — urgency against importance, a different axis pair entirely
We do not know what belongs on the list, or what users are trying to achieveUser Story Mapping — build the list before ranking it
We have thirty ideas, similar in size, and need a rough order by FridayICE Scoring — you are in the right place

What Is It?

Score each idea one to ten on how much it would move the thing you care about, one to ten on how sure you are it will work, and one to ten on how easy it is. Multiply. Sort by the result. That is the whole method, and it takes about a minute per idea once a team has agreed what the scales mean.

It came out of growth experimentation, where the list is long, the items are small and similar in size, and the cost of getting the order slightly wrong is low. Those conditions matter. ICE is built for triage, not for defending a decision.

The multiplication is what makes it useful and what makes it fragile. Multiplying means a weak score anywhere drags the whole thing down, which is usually the behavior you want: an enormous idea nobody believes in should not outrank a decent idea that will certainly work. But it also means small changes to any input move the ranking, and all three inputs are guesses.

Confidence is the factor that does the real work, and the one most often filled in carelessly. It is the only one that asks about evidence rather than about the idea, so it is where a team's actual knowledge enters the calculation. Scored honestly, it demotes the appealing ideas nobody has checked. Scored as enthusiasm, ICE becomes a ranking of how much people like things.

Worth saying plainly: the number is not a measurement. It is a way of making three separate judgments explicit and combining them the same way for every item, which is a real improvement on arguing about a list. Treat the output as a starting order to discuss, not a result.

The ICE formula shown as three input boxes multiplied together into one result: Impact scored one to ten, times Confidence one to ten, times Ease one to ten, equals the ICE score
Three estimates go in and one number comes out, which is the method's convenience and its risk. The output looks far more settled than any of its inputs

Quick Reference

Complexity
Low (2/10)
Time to Decision
1-2 hours
Data Required
Low
Team Size
2-8
Objectivity
Low-Medium
Learning Curve
Minutes

The factors

Scoring each factor honestly

Most of what goes wrong with ICE happens at the point of scoring, not in the arithmetic. Each factor asks something different and fails differently.

What each factor asks, and how it gets scored badly.
FactorWhat it should askHow it goes wrong
Impact
1-10
If this works exactly as hoped, how much does it move the one outcome we have named?Scored against no named outcome, so every idea is impactful in its own terms. Name the metric before scoring anything.
Confidence
1-10
What evidence do we have that it will work? A prior test, a similar case, research, or nothing.Scored as enthusiasm. This is the factor that carries the team's knowledge, and treating it as a mood turns the whole score into a popularity ranking.
Ease
1-10
How little work is it? Ten is easiest.Scored the wrong way round by anyone used to Effort in RICE, which inverts the ranking silently for those rows.

Anchor the scale before anyone scores

The commonest symptom of a useless ICE session is that every score lands between five and eight, so the products cluster and the ranking is decided by rounding. The fix is to pick the highest and lowest item on each factor first, as a group, and score everything else relative to those two anchors. It takes ten minutes and it is the difference between a spread and a smear.

Second habit worth keeping: score independently, then compare, and discuss only the items where people are far apart. A four-point gap on Confidence is not noise to be averaged away. It means two people disagree about what the evidence says, and that disagreement is more valuable than any number the session produces.

Choosing

ICE, RICE, and the ICED question

ICE sits in a family of scoring methods that differ in what they count and how they combine it. Choosing between them is mostly about how much you will have to defend the answer.

The scoring family, and what each is for.
MethodFormulaChoose it when
ICEImpact × Confidence × EaseThe list is long, items are similar in size, and a rough order this afternoon beats a defensible one next week.
RICE(Reach × Impact × Confidence) ÷ EffortReach differs a lot between items, or somebody will ask why their idea was cut. Slower, and the extra factor is the one ICE most often misses.
Value vs EffortNo formula; a two-by-twoYou want the trade-off visible in a picture rather than compressed into a number.
ICEDDisputed — see belowEstablish what somebody means by it first. There is no settled definition, and at least two unrelated ones are in circulation.

On ICED, and why this page is about ICE

This page previously described ICED as Impact, Confidence, Ease and Delight, attributed to the product management community around 2010. We went looking for a primary source and could not find one. The one published meaning of ICED we could trace is entirely different: Infrequency, Control, Engagement and Distinctiveness, a framework for growing products that people use rarely.

So the honest position is this. Adding a delight factor to a scoring model is a reasonable idea and some teams do it. But we cannot point you at who proposed it, when, or what the agreed scales are, and a framework reference that cannot do that is not much use. If you want customer delight in a prioritization decision, the Kano Model researches it rather than estimating it, and ICE handles the rest. If someone hands you ICED, ask which four words they mean.

Core Features

  • Three factors, one scale: all scored one to ten, so the product means something
  • Multiplication, not addition: a weak score anywhere pulls the whole thing down
  • Ease, not effort: higher is easier, the opposite of RICE's denominator
  • Anchored scales: highest and lowest named before anything else is scored
  • Independent scoring: compared afterward, with the disagreements discussed
  • A named outcome: Impact is meaningless until you say impact on what

Worked example

Twenty-two ideas, and four disagreements worth more than the ranking

An illustrative composite. A subscription meal-kit business in Santiago, Chile, with a growth team of five. Twenty-two experiment ideas had accumulated and the weekly prioritization meeting had become an argument. They scored the list with ICE.

What the first pass produced, and what changed on the second.
StageWhat it showed
First passEvery item scored between 5 and 8 on all three factors. Scores ran from 125 to 448 and the top eight were within 40 points of each other, which is well inside the noise of a guess. The ranking was decided by rounding.
Anchoring the scalesThe group named the single highest-impact and lowest-impact idea on the list before rescoring, and did the same for Ease. Scores then ran from 8 to 720, and the top three separated clearly.
Where they disagreedConfidence was scored independently. On four of the 22 items the spread between team members was five points or more. Every one of those four turned out to rest on a belief about customer behavior that nobody had tested.
What the score missedThe top-ranked item affected only subscribers already on the annual plan, about 4% of the base. Nothing in ICE asks how many people an idea touches. Adding Reach dropped it to seventh.
What they ranThe four high-disagreement items were converted into cheap tests of the underlying belief rather than full experiments. Two of the beliefs turned out to be wrong, which removed six ideas from the backlog entirely.

The disagreements were worth more than the ranking

The ordered list was useful and unremarkable. The finding was the four items where Confidence scores were five points apart. Each marked a place where the team held an untested assumption and did not know they disagreed about it — and two of those assumptions were false, which cleared six ideas off the list without running any of them.

Note the other lesson. The idea ICE ranked first reached 4% of customers, because reach is not in the formula. That is not a flaw in ICE so much as its stated scope, and it is exactly the condition under which RICE is worth the extra time. When the items on your list differ wildly in how many people they touch, ICE will mislead you confidently.

When to Use

  • A long list of similar-sized ideas needs a rough order quickly
  • Growth or product experiments, which is what the method was built for
  • Items reach broadly similar numbers of people
  • The cost of getting the order slightly wrong is low and reversible
  • A meeting keeps re-arguing the same list and needs a shared method
  • As a first pass before a more defensible scoring exercise

When NOT to Use

  • Reach varies enormously between items, which ICE cannot see
  • The decision will be challenged and has to be defended in detail
  • Items differ hugely in size, so Ease stops being comparable
  • Dependencies matter, since nothing in the score represents them
  • The list itself is wrong, and the real work is finding what is missing
  • A single large irreversible commitment, where a rough score is no help

Choosing

How the scoring goes wrong

The arithmetic almost never fails. The scoring does, and the single number hides it well.

The recurring failure modes and their remedies.
Failure modeWhat it looks likeWhat to do instead
Everything scores 5 to 8Products cluster, the top items are within noise of each other, and rounding decidesAnchor each scale to the highest and lowest item on the list before scoring the rest.
Confidence as enthusiasmThe ideas people like score high, and the score becomes a popularity rankingMake Confidence cite evidence: a prior test, a similar case, research, or nothing.
Ease scored as effortHard items score high because someone inverted the scale, silently reversing those rowsWrite the direction at the top of the sheet. Ten is easiest.
Averaging the disagreementsA four-point Confidence gap becomes a middling number and the disagreement disappearsScore independently. Discuss only where the spread is wide; that is where the value is.
Reach ignoredA high-scoring idea that affects a tiny fraction of users beats a modest one affecting everybodySwitch to RICE when reach varies. This is the specific gap ICE leaves.
The score becomes the decisionThe ranking is executed rather than discussed, and nobody revisits a bad estimateTreat it as an agenda for a conversation. Re-score after anything is learned.
Scoring a list nobody questionedAn efficient ranking of the wrong twenty ideasAsk what is missing before scoring. A score cannot surface an absent option.

Sourced

Evidence, and how to cite it

ICE is attributed to Sean Ellis, and comes from growth experimentation.

Ellis, who coined the term growth hacking, developed the model to prioritize a continuous flow of growth experiments, where the value of a fast rough order outweighs the value of a precise one. It spread through the growth community from around 2010 and was later adopted for general feature prioritization. There is no canonical paper: it is a practitioner method popularized through writing and practice, and worth citing as such rather than as a research result.

Attributed to Sean Ellis; disseminated through growth-hacking practice and subsequently adopted in product management tooling.

We could not find an independent source for ICED as Impact, Confidence, Ease and Delight.

This page previously carried that definition, attributed to the product management community around 2010. Searching for a primary source returned nothing beyond this site's own earlier page. The one traceable published use of ICED in product management is unrelated: Infrequency, Control, Engagement and Distinctiveness, proposed for growing products that people use infrequently. Adding a delight factor to a scoring model is a sensible thing to do and some teams do it, but we cannot tell you who proposed this version, when, or what the agreed scales are. The page has been rewritten around ICE, which is attributable.

Search conducted August 2026; no primary source located for the four-factor Delight variant. For customer delight assessed rather than estimated, see the Kano Model.

Multiplying three estimates produces a number more precise than its inputs.

Each factor is a judgment on a coarse scale, and the product of three such judgments inherits all their uncertainty while looking like a measurement. A score of 384 against 360 is not a meaningful difference, but it sorts as one. This is a general property of composite scores rather than a flaw peculiar to ICE, and the practical consequence is the same: treat adjacent scores as ties and spend the argument on the gaps that are large.

General property of multiplicative composite indices; see also the discussion of estimate reliability on the RICE page.

There is no evidence that scoring frameworks produce better decisions.

No controlled comparison shows that teams using ICE, RICE or any similar model outperform teams that discuss and decide. What these methods reliably do is make the basis of a decision explicit and consistent across items, which shortens arguments and leaves a record of what was assumed. That is worth having. It is a different claim from the one usually made for them, and the honest reason to use one is the conversation it forces rather than the ranking it outputs.

Assessment of the product management literature as of 2026; no controlled outcome studies of prioritization scoring are available.

How to cite it.

Harvard: there is no single canonical publication. Cite as: Ellis, S. (c. 2010) The ICE scoring model, developed for growth experiment prioritization.
For RICE, cite Intercom's published account of the model. For delight assessed with customers, cite Kano, N. (1984). Do not cite ICED without first establishing which definition your source is using.

Key Strengths

  • Fast: about a minute an item once the scales are agreed
  • Forces three separate judgments: which a discussion tends to blur together
  • Multiplication punishes weak links: an unbelieved big idea does not float to the top
  • Surfaces disagreement: the spread on Confidence is the most useful output
  • Nothing to learn: a team can use it correctly the first time

Key Weaknesses

  • Blind to reach: the specific gap RICE exists to fill
  • Scores cluster: without anchoring, everything lands mid-scale
  • Looks more precise than it is: three guesses, one confident-looking number
  • Ease is easy to invert: silently reversing part of the ranking
  • Cannot see what is missing: it ranks the list it is given

Sequencing

What to run before and after

A score orders a list. It does not produce one, and it does not tell you whether the order was right.

Before

Build the list, and name the outcome Impact refers to

Impact is meaningless until you have said impact on what. And a ranking cannot surface an option nobody wrote down, so it is worth checking what is missing before scoring anything.

During

Score independently, then argue about the gaps

The ranking is the lesser output. Where Confidence scores are far apart, the team is disagreeing about evidence, and turning that into a cheap test is usually worth more than running the top-ranked item.

After

Fix the scope boundary, and re-score when you learn something

A ranked list still needs a line drawn across it for a given release. And every estimate in the score was a guess that new information should update.

Common questions

ICE scoring: quick answers

What is ICE scoring?

A prioritization method that scores each idea on Impact, Confidence and Ease, typically one to ten, and multiplies the three into a single number you can sort on. It was popularized by Sean Ellis for prioritizing growth experiments, and is now used more broadly for features and projects.

What does each ICE factor mean?

Impact is how much this moves the outcome you care about, if it works. Confidence is how sure you are that it will work, which should rest on evidence rather than enthusiasm. Ease is how little effort it takes, scored so that higher is easier. Getting Ease the right way round matters: it is multiplied, not divided.

What is the difference between ICE and RICE?

RICE adds Reach, how many people are affected, and divides by Effort instead of multiplying by Ease. That makes RICE slower and more defensible, and better when you have to justify what you cut. ICE is faster and rougher, and better for triaging a long list of small experiments where reach is roughly similar across them.

What does ICED mean?

It depends who is using it, and we could not find a settled definition. One published usage is Infrequency, Control, Engagement and Distinctiveness, a framework for growing products people use rarely, which is unrelated to scoring a backlog. Some sources describe a four-factor variant of ICE that adds Delight, but we found no primary source for it. If you have been given ICED as a method, ask which of these is meant before using it.

Is a higher or lower ICE score better?

Higher. All three factors are scored so that more is better, including Ease, where ten means easiest rather than hardest. This trips people up because Effort in RICE runs the other way, and a team that mixes the two conventions in one spreadsheet will produce a ranking that is exactly backwards for some rows.

Why do ICE scores end up all the same?

Because people score toward the middle. If everything lands between five and eight on each factor, the products cluster and the ranking is decided by rounding. Force the range: make somebody name the highest and lowest item on each factor first, and score the rest relative to those anchors.

Should ICE scores be averaged across a team?

Averaging hides the thing worth looking at. Where two people score Confidence four apart, that gap is information about a disagreement over evidence, and averaging it to a middling number discards it. Score independently, then discuss only the items where the spread is wide.

How do you cite ICE scoring?

There is no single canonical paper. ICE is attributed to Sean Ellis, who developed it for prioritizing growth experiments and popularized it through his writing and the growth-hacking community from around 2010. For RICE, cite Intercom's published account. Both are practitioner methods rather than research results, and worth citing as such.

Deep Resources