How do I estimate work from a backlog I didn't write?
7 min read · Updated
Estimate from a backlog you didn’t write by splitting it three ways first: what you can size, what needs one question answered before you can size it, and what cannot be sized as written. Quote the three buckets differently, and show the client the split. A single blended number hides exactly the risk that will later cost you.
The instinct is to estimate the work. On an inherited backlog the first job is to estimate the clarity, because clarity is what determines whether your numbers mean anything.
Why not just read the backlog and estimate it?
Because at proposal speed you will read a fraction of it and generalise from that fraction. Four hundred stories, two days to respond, and a team that is also delivering other work: you open the epics, skim thirty stories, form an impression, and price the impression. The impression is drawn from the stories you happened to open, while the estimate has to survive the ones you didn’t.
Sampling fails here in a specific way. It gives you a decent read on the averageand tells you nothing about the tail, and estimates don’t break on average stories. They break on the handful that turn out to be three features wearing one title.
Score the whole backlog before you commit to a date (15 seconds) — what each shot says
- “You didn't write this backlog. You still own the defect.”
- “This one looked ready.” — story PLT-2841, “Users can manage their notification preferences”, with an empty acceptance criteria field.
- “2 sprints later, it shipped.” — the same story, scored 34% and flagged at risk, now linked to defect PLT-3190, “Notification settings silently reset after a password change.”
- “Score your client's whole backlog before you commit to a date.” — a counter runs up past 96,000 stories as score bars fill in across two columns of a backlog.
- “Hand the client a report with your name on it.” — a Backlog Quality Review cover page branded to the agency.
- Closing card: Vindex for Agencies — “Score any backlog. Find the defects before you inherit them.”
The clip is a 15-second version of the argument, and the counter in it is making a point about throughput rather than about any one client: scoring is not something you do to a sample. Vindex has no row limit on an upload, and six-figure backlogs score unattended, so “read everything before quoting” becomes a normal step in a proposal rather than an ambition.
What are the three buckets?
- Sizeable. Clear actor, bounded scope, testable outcome. Estimate these the way you always would; your normal accuracy applies because the input is normal.
- One question away. Broadly clear but missing a specific fact: a volume, a rule, an edge case. These are sizeable after a short clarification pass. Quote them with the clarification included as work, and carry a contingency until the answer arrives.
- Not sizeable as written. No acceptance criteria, unbounded scope, or an outcome no test could confirm. Any number you put on these is fiction with a decimal point.
A quality score sorts the backlog into these buckets far faster than reading does, because the findings map onto them directly: missing acceptance criteria and untestable outcomes land in the third bucket, a single missing constraint lands in the second.
What do I do with the third bucket?
Never fix-price it silently. You have three honest options, and all three beat absorbing the risk:
- Make rewriting them the first milestone. A paid discovery or requirements-hardening phase, priced on its own, with the fixed bid for build quoted afterwards against stories that can actually carry one. This is the strongest position and often the easiest sell, because the evidence for it is concrete.
- Quote them as a range, with the reason attached. “These eleven stories are 40–120 hours depending on answers to the following questions” is a defensible sentence. A single number in the middle of that range is not.
- Carve them out as time-and-materials inside an otherwise fixed-price engagement, named explicitly in the proposal.
How do I present the split without sounding evasive?
Lead with the part you are confident about. “Sixty-one percent of this backlog we can fix-price today, and here is that number” reads as competence; the caveats that follow read as rigour rather than hedging. Reverse the order and the same content reads as an agency trying not to commit.
Then make the unsizeable stories the client’s decision rather than your problem. They can answer the questions, accept a range, or fund a hardening phase. Presented as a choice with a recommendation, it lands as partnership. Presented as a complaint about their documentation, it lands as excuse-making before the work has started — presenting findings to a client covers the framing in depth.
What does this change about the estimate itself?
It gives every number a stated basis. When a client asks why a story is 30 hours rather than 12, the answer stops being experience and becomes evidence: the story scores below the bar on testability, the boundary is undefined, and the range reflects that. Estimates with a documented basis survive procurement, and they survive the mid-project conversation where someone asks how the scope grew.
Once signed, the same scores become your baseline for scope conversations; protecting fixed-bid margins from scope creep picks up from there. If you want to run the whole read as a paid engagement in its own right, see pricing a backlog audit.
Related guides
Pricing models for a productized backlog audit: fixed-fee tiers by backlog size, what to include in each tier, and how to anchor the price against the cost of the defects the audit prevents.
Scope creep on fixed-bid work starts in the requirements, not the change requests. How agencies and consultancies catch vague, oversized stories before signing, and put evidence behind every scope conversation.
