Skip to content
← All writing

You don't need a scoring framework

· 5 min read · by Tan Gravam

The short answer

Most engineering managers reaching for RICE or WSJF do not have a scoring problem; they have a shaping problem. A score compresses judgement into a number, which is fine when the inputs are real — and precision theatre when reach, impact and effort are guesses about a demand nobody has shaped. At commitment scale, two plainer mechanisms do the work the score pretends to: a readiness bar that decides what may be ranked at all, and a visible per-team capacity line you walk the ready demands down until it is spent. Scoring has a legitimate home — ranking hundreds of feature ideas in a roadmap tool — but that is a different layer answering a different question.

The question usually arrives in a tooling frame: should we prioritise with RICE or WSJF? For most engineering managers asking it, the honest answer sits outside the frame: you do not have a scoring problem. You have a shaping problem, and a scoring framework will hide it under two decimal places.

What a score is

A score is judgement compressed into a number, and compression is legitimate when the inputs are real. RICE asks for reach, impact, confidence and effort. For a well-understood demand those can be estimates; for an unshaped one they are guesses wearing units. Score "reporting improvements" and you get a reach nobody counted, an impact nobody defined, and an effort figure the delivering team never saw — multiplied, divided and ranked to a precision the inputs cannot carry. That is precision theatre: the arithmetic is exact and the claim is fiction. Worse, the fiction is load-bearing — a 5.3 beats a 4.9 in a way that ends the conversation, which is the score doing exactly what it was brought in to do.

Where scoring genuinely belongs

Be fair to the frameworks: they were not invented for this moment. A product organisation triaging hundreds of feature ideas against thousands of users needs aggregate instruments, because at that volume you cannot have a conversation about each idea. That is the roadmap-tool layer — feedback in, scored and ranked ideas out — and scoring is honest work there. How that layer differs from this one is its own page, but the short version: roadmap scoring chooses what to explore. It was never designed to decide what a specific team promises a specific quarter, and it performs badly when asked to.

What to use instead of a scoring model

Your quarter is not three hundred ideas. It is perhaps twenty demands and the question of which ones to promise. At that scale, two plainer mechanisms do the work the score pretends to do.

A readiness bar decides what may be ranked at all. A demand enters the planning conversation when it is decision-ready — problem stated apart from the solution, an outcome, a named owner, a capacity figure from the delivering team, dependencies, unknowns written down. Most of the prioritisation problem dissolves right here, because most of the twenty were never rankable: they were topics, solutions in disguise, or wishes with no owner. Scoring them does not resolve the vagueness; it ranks it.

A visible capacity line does the cutting. Walk the ready demands down each team's honest capacity, in the order the organisation's goals suggest, until the capacity is spent. The line does what the score claimed to do: it forces the trade-offs into the open, one at a time, in a form everyone in the room can check — because "these two don't both fit in the backend team's quarter" is arithmetic, not opinion.

Ordering eight things is a conversation

What survives the bar is perhaps eight demands, and ordering eight things is an argument worth having out loud for an hour — not a formula. The score's real appeal was never accuracy; it was shelter. A number lets the room avoid saying "the CFO's request loses to the retention work, and here is why" — the model decided, so nobody chose. But a prioritisation nobody chose is a prioritisation nobody defends when it is challenged in week five, and it will be challenged in week five.

A worked example

Deniz, a director of engineering at an HR-software company, ran the scored version for a year: two afternoons a quarter, sixty-odd backlog items, a RICE sheet. One quarter's top item was "reporting improvements" — reach 4,000 (the user count of the entire module; nobody had counted the reporting users), impact "high", effort 2, entered by a PM who never asked the team. It scored 5.3 and beat items with real evidence behind them, because a number is believed in proportion to how numeric it looks. It shipped six weeks late, as something its requesters did not recognise.

The next quarter Deniz dropped the sheet. The readiness bar filtered fifty-eight items to nine that were decision-ready or one question away. Walking those down the three teams' capacity fit six. The ordering discussion took forty minutes and produced two hard trade-offs, both made by named people with reasons on the record — so when the week-five challenge came, the answer was a recorded decision, not a spreadsheet cell nobody could reconstruct.

An honest note about tooling

DeliverySheet, the product behind this site, has no scoring framework — deliberately. It holds the step before scoring: making the thing being ranked real. If you genuinely operate at feedback volume — hundreds of ideas, thousands of users — a roadmap tool with scoring on top is a coherent companion: score to choose what to explore, then govern what you promise. What you do not need is a formula standing between twenty demands and a promise. You need a bar and a line.

I'm Tan Gravam. I build DeliverySheet — it takes a vague work request to a clear delivery decision, so the shape, owner, capacity, dependencies and open questions are on the table before anyone commits people or a date.

$189/month per workspace, unlimited members, 7-day free trial. I answer the support email myself.

More on deciding what to commit to

Read the overview: How to decide what to commit to →