The ICE scoring model prioritizes ideas by rating each on three dimensions, impact, confidence, and ease, typically on a scale of one to ten, and multiplying or averaging the scores to produce a ranking. Associated with Sean Ellis and widely used in growth teams, it is deliberately lightweight, designed to order a backlog quickly rather than to produce a defensible estimate of value.
Its components map onto the questions that actually determine whether an idea is worth doing. Impact estimates how much the change would move the target metric if it worked. Confidence estimates how sure the team is that it will work, based on supporting evidence rather than enthusiasm. Ease estimates how little effort is required, which is the inverse of cost and is included because a portfolio of cheap experiments generates more learning per quarter than a small number of expensive ones.
The confidence dimension is what distinguishes the model from simpler impact-versus-effort ranking, and it is the one most often applied carelessly. Confidence should reflect evidence: a hypothesis supported by session recordings, analytics, customer interviews, and a similar result observed previously deserves a high score, while an idea someone found compelling in a meeting does not. Teams that assign confidence based on conviction rather than evidence convert the model into a mechanism for ranking opinions, which is what it was meant to replace.
The framework's weakness is that the scores are subjective and the arithmetic gives them unwarranted authority. Two people can score the same idea very differently, and multiplying three uncertain estimates produces a number with far more apparent precision than the inputs support. A score of 240 is not meaningfully better than one of 210, and treating the ranking as exact is a misuse. The model works as a sorting mechanism to separate obvious priorities from obvious deprioritizations, and to structure discussion about disagreement, rather than as a calculation.
Consistency is what makes it usable in practice. Scoring works reasonably when the same people score in the same session using shared definitions, since relative comparisons hold even if absolute values are arbitrary. It breaks down when different people score at different times without agreed anchors, when impact is scored against different metrics for different ideas, or when scores are assigned by whoever proposed the idea. Simple conventions, such as scoring as a group, defining what each level means with examples, and requiring the evidence behind a confidence score to be stated, resolve most of this.
Comparing frameworks is less useful than choosing one and applying it consistently, since the main benefit of any scoring model is that it makes the reasoning behind prioritization explicit and comparable across ideas. More elaborate frameworks that separate reach from impact, or that incorporate effort as a divisor, add precision that is largely illusory given the quality of the underlying estimates, though they can be worth the additional effort where individual items are expensive enough that the ranking genuinely matters. The failure mode common to all of them is treating the output as a decision rather than as an input, and any framework will produce poor results if the scores are assigned to justify a conclusion already reached.
ICE is best understood as a faster alternative to more elaborate frameworks, appropriate where the cost of each item is low and the number of candidates is high, which describes most experimentation backlogs. In a CRO service programme it is commonly used to order test ideas at each planning session, with the sample size and expected runtime of each test considered alongside the score, since an idea that scores well but requires two months of traffic competes differently from one that resolves in ten days.