top of page

The RICE Prioritization Framework Explained

Shawn West
Apr 8
7 min read

Updated: 2 days ago

RICE — devised by Sean McBride on Intercom's product team in 2016, adding Reach to the older ICE model — gets adopted as a ranking machine and abandoned as false precision, and both reactions miss what it's for. The number is nearly worthless on its own. What earns its keep is the moment two people score the same idea differently — because the four inputs tell you exactly which conversation you owe, and to whom.


A product team scored fourteen ideas for the quarter. One came back with a wildly split result: the PM had it at 80, an engineer had it at 1,440. Same idea, same afternoon, same room.


The instinct is to average them, or to argue about whose judgment is better. Both waste the signal. The whole value of decomposing a priority into four inputs is that a divergence lands in one of them — and each input belongs to a different person, answerable by a different kind of evidence. Find the input, and the argument stops being about opinion.



The formula, and what it's really doing


RICE = (Reach × Impact × Confidence) / Effort

The scales, so two people's numbers mean the same thing: Reach is a count of people or events per time period. Impact uses RICE's ordinal scale — 3 massive, 2 high, 1 medium, 0.5 low, 0.25 minimal. Confidence is a percentage. Effort is person-months, or person-weeks if you keep it consistent across the sheet. An Impact of 2 below means "high," not two of anything.


  • Reach — how many people, in a fixed period

  • Impact — how much per person, on a named metric

  • Confidence — your discount for uncertainty

  • Effort — person-months, from the people who'd build it


Read as a ranking device, this is a machine for turning guesses into a number with a decimal point — which is the version critics are right to dislike. Read as a decomposition, it's doing something else: it splits one unanswerable question ("is this more important?") into four answerable ones, each owned by someone specific.


That's the mechanism worth carrying. A single priority score can only be agreed or disagreed with. Four inputs can be located. When two people differ, exactly one or two of the inputs are usually responsible, and the fix is a different action each time.


The test you can run: take any two people's scores for the same item and diff the four inputs before you look at the totals. If you can't say which input drove the gap, you scored a feeling and decorated it with arithmetic.


The 80 versus 1,440


(Developed example — a simple scenario.)


Here's what the two scoring sheets actually held.


Input

PM

Engineer

Same?

Reach

50 / month

900 / month

No — 18×

Impact

2

2

yes

Confidence

0.8

0.8

yes

Effort

1 person-month

1 person-month

yes

Score

80

1,440

—


Three of four inputs matched exactly. The entire gap was Reach, and Reach is not a matter of judgment — it's a query.


The idea was bulk export in the admin panel. The PM had scored the people who'd asked for it: fifty admins, from support tickets. The engineer had scored the people who hit the underlying endpoint: nine hundred a month, because a scheduled integration several customers had built was calling the single-record export in a loop, and he'd seen it in the logs while investigating a rate-limit alert.


Both numbers were accurate. They were counting different populations, and neither knew the other existed. The engineer's number changed the decision — but more importantly, it surfaced that several customers had built a fragile workaround nobody had documented, which is a finding worth more than the ranking it came from.


The resolution took nine minutes: pull the actual endpoint call volume, agree the population, rescore once. Reach settled at 780 after excluding internal calls.


The test: for your highest-scoring item, ask who exactly did we count, and where did that number come from? A Reach figure with no query behind it is the most common way a RICE sheet is confidently wrong.


Each input is a different person's question


The reason to locate the disagreement is that each input escalates somewhere different. Getting this wrong is why RICE sessions turn into unresolvable debates — the room argues about the score instead of routing the question.


Input

A disagreement here means

Who resolves it

How

Reach

You're counting different populations, or someone has no data

Whoever owns the analytics or the logs

Run the query. This is a fact, not a judgment.

Impact

You disagree about which metric moves, not how much

Product, with the metric named

Name the metric first, then re-score. "Impact 2" with no metric is unscoreable.

Confidence

One of you has evidence the other hasn't seen

Whoever holds the research, tickets, or prior experiment

Put the evidence on the table. If neither has any, that's the finding.

Effort

The two of you are sizing different scopes

The engineers who'd do the work

Ask what each was sizing — the same estimate-spread question that splits stories in INVEST.


Read down the third column: not one of these is resolved by the person running the session. A facilitator who tries to adjudicate is guessing on someone else's behalf, and the score inherits the guess.


Confidence is the input everyone quietly deletes


Most teams set Confidence to 100% for everything, which removes the only term in the formula that represents how much you actually know. What's left ranks strictly by expected value with certainty assumed — which systematically over-promotes ambitious bets.


Honest anchors — a finer-grained variant of Intercom's canonical 100 / 80 / 50% tiers:


  • Data-backed, similar to work we've shipped → 80–100%

  • Users have told us directly → 60–80%

  • We believe users want this → 40–60%

  • We're guessing → 20–40%


The number matters less than the threshold it creates. Below about 50%, the useful question stops being "should we build this?" and becomes "what's the cheapest thing that would raise the confidence?" That's usually an interview, a query, or a prototype — the instruments in requirements elicitation — costing days rather than the months the feature would.


A low-confidence item isn't a rejected item. It's an item whose next step is discovery rather than delivery, and the formula is what makes that visible instead of leaving it as a hunch someone overrides.


The test: count how many items on your sheet have Confidence at 100%. If it's most of them, the column is decorative and your ranking is uncertainty-blind.


When RICE is the wrong instrument


Forcing RICE onto everything produces scores nobody believes, and the damage is that it discredits the tool where it genuinely works.


There is also a structural bias worth naming: effort sits in the denominator, so RICE systematically penalises large-but-necessary work. A platform migration or a compounding investment divides a modest numerator by a large number and never ranks, however right it is. If your sheet has never surfaced an infrastructure item, that is the arithmetic talking, not the backlog. The mirror of the Confidence failure applies too — teams that delete the column lose the hedge, but teams that write 80% against a hunch have laundered a guess into something that looks defensible.


It doesn't fit when:


  • The work is non-negotiable — a regulatory deadline, a security patch under active exploitation. Scoring it is theatre; it happens regardless of where it lands.

  • Reach isn't a population — a strategic bet on a market you're not in yet has no denominator. The number would be invented.

  • The item's whole purpose is to reduce uncertainty — a spike scores terribly by construction, because low Confidence is the point. Judge it on what it would let you decide, not on expected value.

  • You're filling a fixed container — RICE ranks a continuous backlog. If you have a date and a capacity ceiling, you want MoSCoW, which asks whether the release fails without an item rather than how it ranks.


The two pair cleanly rather than competing: RICE produces the order, MoSCoW decides how far down that order this release reaches. And for what you do next week — as opposed to what the team builds this quarter — neither applies; that's the Eisenhower matrix.


RICE ranks the options; it does not say who makes the call. The DACI Decision Framework for Engineering Teams covers that half.


The test: before scoring an item, ask whether a low score would actually stop you doing it. If it wouldn't, take it off the sheet — it's a commitment, not a candidate, and leaving it in distorts everything ranked against it.


What to do at your next scoring session


Change one thing: score independently, then diff the inputs before anyone sees a total.


Have two or three people score the same set alone. Then, for every item where the totals differ by more than roughly 3×, put the four inputs side by side and find the row that caused it. Don't discuss the item — discuss the row, and route it to whoever owns that input. Most gaps close in minutes because they turn out to be a query somebody can just run.


Then check one column: how many items sit at 100% Confidence. Each one is a claim that you have data. For every item where that isn't true, write down the cheapest thing that would raise it, and compare that cost against the build.


The framework's output was never the ranking. It's that a number nobody can argue with productively gets decomposed into four questions that each have an owner and an answer — and the disagreements you surface that way are usually worth more than the order you end up with. The team in the example got a defensible priority. What they actually got was the discovery that customers were looping a single-record endpoint because the feature they needed didn't exist.


Sources


Sean McBride, "RICE: Simple prioritization for product managers," Intercom, 2016 — the original definition of the framework and its Reach / Impact / Confidence / Effort inputs, including the ordinal Impact scale used above.

bottom of page