How do I estimate work from a backlog I didn't write?
7 min read · Updated
Estimate or plan from a backlog you didn’t write by splitting it three ways first: what you can size, what needs one question answered before you can size it, and what cannot be sized as written. Forecast the three buckets differently, and show the sponsor the split. A single blended date or number hides exactly the risk that will later be blamed on the plan.
The instinct is to estimate the work. On an inherited backlog the first job is to estimate the clarity, because clarity is what determines whether your numbers mean anything. This applies equally to the PMO consultant asked for a roadmap, the program lead asked to re-baseline, and the delivery firm asked for a price.
Why not just read the backlog and estimate it?
Because at planning speed you will read a fraction of it and generalise from that fraction. Four hundred stories, two days before the steering committee, and a team that is also delivering other work: you open the epics, skim thirty stories, form an impression, and plan the impression. The impression is drawn from the stories you happened to open, while the plan has to survive the ones you didn’t.
Sampling fails here in a specific way. It gives you a decent read on the averageand tells you nothing about the tail, and plans don’t break on average stories. They break on the handful that turn out to be three features wearing one title.
Score the whole backlog before you commit to a date (15 seconds) — what each shot says
- “You didn't write this backlog. You still own the defect.”
- “This one looked ready.” — story PLT-2841, “Users can manage their notification preferences”, with an empty acceptance criteria field.
- “2 sprints later, it shipped.” — the same story, scored 34% and flagged at risk, now linked to defect PLT-3190, “Notification settings silently reset after a password change.”
- “Score your client's whole backlog before you commit to a date.” — a counter runs up past 96,000 stories as score bars fill in across two columns of a backlog.
- “Hand the client a report with your name on it.” — a Backlog Quality Review cover page branded to the firm.
- Closing card: Vindex for Agencies — “Score any backlog. Find the defects before you inherit them.” (The program has since been renamed Vindex for Consultancies.)
The clip is a 15-second version of the argument, and the counter in it is making a point about throughput rather than about any one client: scoring is not something you do to a sample. Vindex has no row limit on an upload, and six-figure backlogs score unattended, so “read everything before committing” becomes a normal step in a planning cycle rather than an ambition.
What are the three buckets?
- Sizeable. Clear actor, bounded scope, testable outcome. Estimate these the way you always would; your normal accuracy applies because the input is normal.
- One question away. Broadly clear but missing a specific fact: a volume, a rule, an edge case. These are sizeable after a short clarification pass with the product owner. Plan them with the clarification included as work, and carry a contingency until the answer arrives.
- Not sizeable as written. No acceptance criteria, unbounded scope, or an outcome no test could confirm. Any date or number you put on these is fiction with a decimal point.
A quality score sorts the backlog into these buckets far faster than reading does, because the findings map onto them directly: missing acceptance criteria and untestable outcomes land in the third bucket, a single missing constraint lands in the second. The bucket sizes are also a coaching finding in their own right: a backlog that is 40% unsizeable has a refinement problem, not a planning problem.
What do I do with the third bucket?
Never commit to it silently. You have three honest options, and all three beat absorbing the risk into the plan:
- Make rewriting them the first milestone. A requirements-hardening phase, run as workshops with the product owner and sized on its own, with the plan for the rest committed afterwards against stories that can actually carry one. This is the strongest position and often the easiest sell, because the evidence for it is concrete and the workshop is the coaching the team needed anyway.
- Forecast them as a range, with the reason attached. “These eleven stories are 40–120 hours depending on answers to the following questions” is a defensible sentence. A single number in the middle of that range is not.
- Carve them out of the committed plan explicitly, as a named uncertainty in the roadmap or, for delivery firms, as time-and-materials inside an otherwise fixed-price engagement.
How do I present the split without sounding evasive?
Lead with the part you are confident about. “Sixty-one percent of this backlog we can commit to today, and here is that plan” reads as competence; the caveats that follow read as rigour rather than hedging. Reverse the order and the same content reads as a consultant trying not to commit.
Then make the unsizeable stories the sponsor’s decision rather than your problem. They can have the product owner answer the questions, accept a range, or fund a hardening phase. Presented as a choice with a recommendation, it lands as partnership. Presented as a complaint about the team’s documentation, it lands as excuse-making before the work has started — presenting findings to a client covers the framing in depth.
What does this change about the plan itself?
It gives every number a stated basis. When a sponsor asks why a story is 30 hours rather than 12, or why the date moved, the answer stops being experience and becomes evidence: the story scores below the bar on testability, the boundary is undefined, and the range reflects that. Plans with a documented basis survive the steering committee, and they survive the mid-project conversation where someone asks how the scope grew.
Once committed, the same scores become your baseline for scope conversations and for the re-score that shows the hardening worked. For delivery firms, protecting fixed-bid margins from scope creep picks up from signature. If you want to run the whole read as a paid engagement in its own right, see pricing a backlog audit.
Related guides
Pricing a backlog assessment as a productized service for consultancies and training firms: fixed-fee tiers by backlog size, a diagnostic, a workshop tier, and a measured-engagement tier, anchored against the spend the assessment de-risks and proves.
The delivery lane: scope creep on fixed-bid work starts in the requirements, not the change requests. How agencies, dev shops, and consultancies that take fixed-fee work catch vague, oversized stories before signing, and put evidence behind every scope conversation.
