Rank Your CRO Backlog When ICE, PIE, PXL and RICE Disagree

Rank Your CRO Backlog When ICE, PIE, PXL and RICE Disagree

Only about one third of the ideas tested on Microsoft's experimentation platform improved the metric they were designed to improve, and Netflix has put its own hit rate closer to one in ten (Kohavi and Longbotham, 2015). If most of your backlog is wrong, the order you test it in is most of the return you will get. This guide is for CRO leads and product owners who already have more test ideas than traffic. We score the same three ideas with ICE, PIE, PXL and RICE, show why the four rank orders disagree, give a rule for choosing between them, and add the gate all four are missing: whether a test can reach significance before the quarter ends.

What each framework is actually betting on

All four produce a number, and none measures the same thing. Each is less a scoring system than a compressed opinion about what makes a test worth running.

FrameworkInputsFormulaOriginThe bet it encodes
ICEImpact, Confidence, Ease, each 1 to 10I × C × ESean Ellis, growth practiceExperienced judgement is the cheapest signal
PIEPotential, Importance, Ease, each 1 to 10Average of the threeChris Goward, WiderFunnel (now Conversion)Where you test matters more than what you change
PXL~9 mostly binary questions, some weighted 0 or 2Sum, out of a fixed maximumPeep Laja, CXL, 27 Sep 2016Evidence behind an idea beats confidence in it
RICEReach (people per period), Impact (0.25 to 3), Confidence (50/80/100%), Effort (person-months)(R × I × C) / ESean McBride, Intercom, 5 Jan 2018Score impact per unit of work

Two corrections before you adopt one. PXL was written as an attack on the other CRO frameworks: Laja's argument is that if you could reliably guess potential or confidence in advance, you would not need a test. And PIE's canonical page defines the criteria but publishes no scale and no formula, so every team invents its own. The widerfunnel.com domain now redirects to conversion.com, so the framework most CRO decks cite has no maintained primary specification.

The same backlog, scored four ways

Three candidate tests on a mid-size retailer. Figures are illustrative, not client data.

  • A. Instalment (taksit) breakdown inside the product page price block. 180,000 PDP viewers a month, supported by analytics and voice-of-customer tickets.
  • B. Remove forced account creation in checkout. 12,000 checkout starts a month, supported by usability sessions, recordings, analytics and support contacts.
  • C. Rewrite the homepage hero headline. 220,000 homepage sessions a month, no supporting research, proposed in a workshop.
TestRICEICEPIEPXL (of 12)
A: installment breakdown144,000 (1st)294 (2nd)7.0 (2nd)10 (1st)
B: no forced account18,000 (3rd)405 (1st)6.3 (3rd)9 (2nd)
C: hero headline55,000 (2nd)144 (3rd)7.3 (1st)6 (3rd)

Every test wins under at least one framework. The disagreement is the frameworks doing their jobs:

  • RICE puts A first because Reach is a raw headcount. It will bury checkout and account work on any normal funnel.
  • ICE puts B first because a senior practitioner scores forced account creation Impact 9, Confidence 9, and ICE multiplies opinion by opinion. Change the scorer, change the roadmap.
  • PIE puts C first because the homepage has the most valuable traffic and copy is trivial to ship. PIE was built to choose pages; asked to rank changes, it rewards cheap edits to important templates.
  • PXL puts C last because C scores zero on all four evidence questions. It is the only framework that penalises a hunch structurally instead of trusting the scorer to be honest about Confidence.

Which one to use: pick by your binding constraint

Binding constraintUseBecause
No research inputs yet, first ten testsICE, with written anchorsWorks on judgement alone. Treat the output as a queue, not a ranking.
The team keeps testing opinions and losingPXLBinary evidence questions are hard to inflate.
Engineering capacity, not trafficRICEThe only one with a denominator, so the only one that optimises throughput.
Choosing which page or template to work onPIEBuilt for that question.
TrafficNone aloneAll four are blind to whether a test can finish. Add the gate below.

The gate all four are missing: can this test finish?

Reach is not detectability. Check feasibility before scoring, using the normal approximation for a two-sided test at 95% significance and 80% power:

n per variant ≈ 15.7 × p × (1 - p) / d²

where p is the baseline rate of the metric you are moving and d is the absolute lift you want to detect. Applied to the backlog above:

TestMetric and baselineMDEPer variantWeeks
APDP viewer to purchase, 2.0%5% relative (0.1 pp)~307,700~14
APDP viewer to purchase, 2.0%10% relative (0.2 pp)~76,900~4
BCheckout start to purchase, 55%5% relative (2.75 pp)~5,140~4

The test RICE ranked first cannot be read for a quarter at the sensitivity most teams assume, while the test RICE ranked last finishes in a month. A sits on a 2% base and B on a 55% base, and sample size is driven by p(1-p)/d², not by how many people saw the page. A also becomes feasible the moment you declare 10% the smallest lift worth acting on, which makes your minimum detectable effect a prioritisation decision rather than a statistics detail. Decide it in the same meeting as the scoring, and pick one calculator for everything, since different variance approximations do not agree exactly: Evan Miller's and the Switas A/B test calculator are both fine.

Feasibility buckets

  • Two weeks or less. Green. Score it and run it.
  • Two to four weeks. Amber. Run it alone, covering at least two full weekly cycles.
  • Four to eight weeks. Red for a standard A/B test. Raise the MDE, move to a higher-base metric, or use a sequential design with a pre-registered stopping rule.
  • Over eight weeks. Do not A/B test it. Ship on judgement with guardrail metrics, or run a qualitative study.

This gate matters more here than the English-language CRO literature assumes. Turkey's e-commerce volume reached 4.567 trillion lira across roughly 634,000 businesses in 2025 (Ministry of Trade, 12 May 2026), an average of about 7.2 million lira per business per year, and the distribution is skewed toward a handful of marketplaces, so the typical merchant sits well below it. A site doing a few hundred orders a month cannot power a checkout test at any useful MDE. For those sites the honest output of prioritisation is a research and measurement roadmap, not a test roadmap.

Making the scores mean something

  1. Write anchors before scoring. Impact 3 in RICE must mean something in your currency, for example "moves purchase conversion by more than 5% relative". Without anchors, scores drift within a single session.
  2. Cap Confidence by evidence tier. No research caps at 50%, analytics only at 80%, analytics plus qualitative plus a related prior win at 100%. This imports PXL's discipline into RICE.
  3. Score blind, reconcile the widest spreads first. Three people scoring independently, then discussing only items that differ by three points or more, takes about forty minutes for thirty items.
  4. Record predicted score against actual result. After ten readouts, check whether Impact scores correlate with observed effects at all. If not, the scale is decoration and you should switch to PXL.
  5. Never compare scores across frameworks. A RICE score of 144,000 and a PXL score of 10 have no relationship.

Where this breaks down

These frameworks rank ideas, they do not improve them. A weak hypothesis scoring 11 out of 12 on PXL is still weak, and no model will tell you the real problem sits upstream of the page you are testing. Specific limits:

  • Effort is the least reliable input and nobody re-scores it. RICE's denominator is the number most likely to be wrong by a factor of two, and that moves rank order more than a wrong Impact score.
  • None of the four models interaction. Run three tests on one funnel at once and the scores told you nothing about interference.
  • Reach in people per month breaks for seasonal businesses. Travel and education should score Reach over a full seasonal cycle, or against the same period last year.
  • Dependencies are invisible. A score shared across a test backlog and a product roadmap also flatters whichever side has more traffic.

[INTERNAL DATA NEEDED: Switas win rate by pillar over the last 24 months, to replace the industry hit-rate figures with first-party numbers.]

[INTERNAL DATA NEEDED: an anonymised client backlog scored under two frameworks, to replace the illustrative example.]

Frequently asked questions

Is RICE better than ICE for CRO work?

Not inherently. RICE adds Reach and puts Effort in the denominator, which helps when delivery capacity is the bottleneck, but it biases toward high-traffic pages. Use RICE when engineering time is scarce, ICE when it is not.

What score is high enough to run a test?

None of the four has an absolute threshold. The scores are ordinal: they tell you what to do next, not whether anything is worth doing. The absolute gate is feasibility. If a test cannot reach significance in eight weeks, its score is irrelevant.

Can we combine two frameworks?

Combining outputs is meaningless because the scales are unrelated. Combining inputs is not. Capping RICE's Confidence with PXL's evidence questions is a well-behaved hybrid and the one we use most often.

Does PXL work for non-e-commerce sites?

The evidence questions transfer to any product. The layout questions assume a marketing page and are weak for logged-in flows. Replace those with questions about how often users meet the element in the task flow.

Should low-traffic sites use these frameworks at all?

Yes, for sequencing work, but stop calling the output a test roadmap. On a site that cannot power tests, the same exercise orders UX audit findings, research questions and build work.

Who should be in the scoring session?

Three to five people who see different evidence: analytics, research or support, and whoever owns delivery. Scoring alone produces confident nonsense; scoring with ten people produces averages that rank nothing.

Try this on your own backlog

Take the ten items at the top of your test list, run each through the feasibility formula above, and count how many survive at the sensitivity you have been assuming. If more than half do not, ranking was never your problem. Our CRO and experimentation team runs this scoring and feasibility pass as the first step of an engagement.

Sources


Inci Dindar
Written by

Inci Dindar

With a background in Software Development, she led design, UX, and product teams across fintech, media, e-commerce, and travel. For the past three years she has consulted full-time on UX and design — including as an external UX consultant for a global management consultancy helping brands and startups with audits, conversion optimization, and design processes.


Related Articles

Switas As Seen On

Magnify: Scaling Influencer Marketing with Engin Yurtdakul

Check Out Our Microsoft Clarity Case Study

We highlighted Microsoft Clarity as a product built with practical, real-world use cases in mind by real product people who understand the challenges companies like Switas face. Features such as rage clicks and JavaScript error tracking proved invaluable in identifying user frustrations and technical issues, enabling targeted improvements that directly impacted user experience and conversion rates.