A/B Testing
The ICE framework is broken. Here's what to use instead for A/B test prioritisation
Last updated: [Updated Date]
If you're still scoring your A/B test backlog with ICE, your test plan isn't a prioritisation system. It's a popularity contest where the loudest hypothesis owner wins. You're not running CRO; you're refereeing meetings.
That's the uncomfortable thing every CRO team eventually notices. Impact, Confidence, Ease, score each out of 10, multiply, prioritise descending. The appeal is obvious: a clean number you can defend in meetings. The problem is that all three inputs are subjective, and in practice the person who presents the hypothesis also scores it.
I've sat through dozens of ICE-scored test backlogs where every advocate scored their own hypothesis 9/10/9 and nothing was ever deprioritised. The maths look like science. The behaviour underneath is closer to a school sports day where everyone gets a participation rosette.
Where ICE actually came from (and why it stopped working)
ICE was popularised by Sean Ellis through GrowthHackers in the early 2010s as a lightweight scoring rubric for growth experiments. Ellis was solving a specific problem: scrappy growth teams needed something simpler than a full PMF analysis to triage their idea backlog. ICE did that job well.
It also spawned two refinements that most CRO teams haven't caught up with. WiderFunnel's Chris Goward built PIE (Potential, Importance, Ease) to push prioritisation toward pages that drive business outcomes, not just any test that catches a hypothesis owner's eye. ConversionXL (now CXL) built PXL, a 14-question evidence-weighted scorecard that grades each input objectively against research artefacts rather than gut feel.
Sean Ellis built ICE for 2012 growth teams that had no research function. Most CRO teams in 2026 do have a research function. They're scoring like it's still 2012.
That's the lineage. The fact that PIE and PXL exist tells you the industry has known about ICE's failure mode for over a decade. Most agencies just haven't updated the script.
The three-part framework that replaces ICE
This is the prioritisation core of OperatorAI (GoGoChimp's CRO methodology, distinct from OpenAI's Operator agent product). Three inputs, weighted by evidence, not advocacy. We've run it across every client engagement since 2019.
How do you score evidence weight in practice?
Score 0 to 5 on what you actually know. A session recording showing 30% of users abandon on this step gets a 4. A past test on a similar page that won at 9% lift gets a 5. Someone's hunch gets a 0.
No hypothesis enters the backlog with fewer than 2 points of evidence weight. That's not a guideline. That's the bouncer at the door.
The evidence-weight axis is the spiritual descendant of CXL's PXL 14-question scorecard. We collapsed it into a single 0-5 because most engagements don't have time for 14 questions per hypothesis. Most teams need the bouncer more than they need the rubric.
What's a reasonable ceiling estimate?
If this variant wins, what's the realistic upper bound on the lift? Base it on comparable tests, industry benchmarks, or Bayesian priors from prior experiments. A 2% ceiling test on high-traffic is worth more than a 20% ceiling test on a low-traffic page.
The trap most teams fall into here: confusing ceiling estimate with hope. The ceiling isn't what the advocate wants. It's the top end of the range observed in similar tests across similar pages. Nielsen Norman Group's guidance on A/B testing is helpful here. Their work on iterative testing makes the point that small improvements compound, and reaching for unrealistic ceilings starves the backlog of the small wins that add up.
How long is too long for a test runtime?
Run the sample-size maths before prioritising. Use Evan Miller's sample-size calculator. If your traffic produces significance in 9 days, ship it. If it requires 47 days, either increase traffic, accept the opportunity cost, or deprioritise.
The reason runtime matters as a third axis: every day a test is running is a day every other hypothesis isn't being tested. Opportunity cost is the third dimension ICE flat-out ignores.
A worked example: scoring three real hypotheses
Imagine you're a Shopify operator with three competing hypotheses in your backlog this month. Here's how ICE would have scored them, and how the new framework actually prioritises them.
Hypothesis A: change the checkout button colour from blue to orange.
Hypothesis B: rewrite the hero headline to lead with the £-figure on average order saving.
Hypothesis C: add a free-shipping threshold progress bar to the cart drawer.
| Hypothesis | ICE score (advocacy) | Evidence weight | Ceiling estimate | Runtime at current traffic | Verdict |
|---|---|---|---|---|---|
| A. Button colour | 9 × 9 × 9 = 729 | 0 (no research) | 1-2% | 21 days | Rejected at door |
| B. Hero headline rewrite | 8 × 7 × 6 = 336 | 4 (session recordings show 41% bounce on hero) | 8-12% | 14 days | Ship first |
| C. Free-shipping progress bar | 6 × 8 × 7 = 336 | 3 (Baymard research on cart friction) | 3-5% | 11 days | Ship second |
ICE would have ranked them A > B = C, sending the team to test button colour first. The framework ranks them B > C > A, rejecting A outright. Evidence weight does the bouncer work; ceiling and runtime decide the order.
The Baymard reference matters. Baymard's cart-abandonment meta-analysis across 50 studies found extra costs (shipping, taxes, fees) drive 48% of abandonment. That's the research artefact that turns hypothesis C from a hunch into a 3-point evidence weight. Without that lookup, it would have stayed at 0.
How this changes the backlog
Tests with no evidence weight go to the bottom regardless of how confident the advocate feels. Tests with 30-day-plus runtimes get staged later than faster tests of similar ceiling. The button-colour hypothesis (which somehow always makes it into a backlog) gets rejected before it consumes a single visitor.
Advocacy stops winning. Evidence wins.
The downstream effect across our portfolio: the win rate per test rises because the no-hope hypotheses get filtered upstream. Build Grow Scale's 2026 review of 347 stores found expert-guided AI testing delivered 28-34% lifts versus 4-7% from DIY tools. A lot of that 6× delta is just prioritisation: testing the right things in the right order.
When ICE is still useful
This isn't a religious argument. ICE works fine in two contexts.
First, when your backlog has under 10 hypotheses and you're still building the testing muscle inside an organisation that's never tested before. ICE is a step up from "the highest-paid person's opinion wins". For a CRO programme month one, it does honest work.
Second, for non-CRO teams that need a shared language. A product team scoring feature requests for the next sprint can use ICE without setting off alarm bells. The stakes are lower; the framework's looseness becomes a feature.
The line: when test results start informing revenue forecasts, when you're shipping more than two experiments a month, when the cost of a wasted runtime exceeds the cost of an extra hour scoring evidence, switch frameworks. CXL's CXL Institute training teaches PXL as the default. We use the three-input variant above. Both beat ICE once you cross that volume threshold.
Where this fits in the OperatorAI methodology
This article sits under The Evidence Stack, one of the three named frameworks inside our OperatorAI methodology. GoGoChimp's four-layer testing discipline: operator-set hypothesis, sample-size discipline, The 99 Rule, and failure-as-information.
For where this work sits in our operating-model maturity classification, see The OperatorAI Maturity Model, the five-tier framework from Ad-hoc through Operator-Led.
Want us to do this for your site?
Book a free AI audit. 15 minutes. We’ll show you three things your site is missing and what we’d test first.
Book my free AI audit →



