How do I audit a backlog when the project is already slipping?
7 min read · Updated
When the project is already slipping, audit the backlog that remains, not the sprints that went wrong. Score what is left, split the slip into the part that traces to requirement quality and the parts that don’t, and put all of it into a single recovery plan. A backwards-looking assessment produces an argument; a forwards-looking one produces a schedule.
This is the hardest assessment a consultant runs, because the room is already taking sides. The sponsor wants to know why the date moved, the team feels accused, any vendor is defending its estimates, and you have been brought in, or were already embedded as the coach, to say something true that all three will accept. Anything you say about the requirements can be heard as taking the vendor’s side; anything about delivery, as taking the team’s.

Why score the remaining backlog first?
Because it is the only part still under anyone’s control, and because it changes the meeting. Scoring delivered work invites a forensic argument about decisions taken months ago, with every party motivated to remember them differently. Scoring the remaining two hundred stories asks a different question: given what we now know about how these requirements are written, what will the next three months actually cost, and what has to change in how stories reach a sprint?
That question has an answer, the answer is measurable by a neutral engine rather than by you, and nobody has to lose face to accept it.
How do I separate requirement quality from delivery?
Deliberately and in public, because a rescue readout that finds one party entirely at fault has no credibility. Do the split explicitly:
- Traceable to requirement quality. Rework on stories that scored below the bar, defects from behaviour that was never specified, estimates that moved after a clarification. Evidence, per story, from the scoring engine.
- Traceable to delivery. Estimation misses, staffing gaps, environment problems, decisions the team or its vendor got wrong. Name these with the same precision if you want the requirements finding to be believed.
- Neither. Third-party dependencies, access delays, reorganizations, market changes. Worth listing so the total adds up.
A distribution across three named causes reads as analysis. A single cause that happens to favour whoever hired you reads as advocacy, even when it is true.
How much of the delivered work should I score?
Enough to establish that the pattern is real, and no more. If ten of the twelve defects in production came from stories that scored in the bottom quartile, that is a finding worth two slides — it demonstrates the mechanism rather than asserting it. Beyond that, more historical scoring buys nothing except a longer argument.
The full causal chain from a low score to a shipped defect is worked through in what a low-scoring story actually costs, which is a useful thing to send ahead of a rescue readout so the mechanism is understood before the numbers about their project appear.
What does the recovery plan contain?
- A quality gate on the remaining backlog. No story enters a sprint below an agreed score. This is the single change that stops the slip compounding, it is cheap, and it is a ways-of-working change rather than a staffing one, which is usually what the sponsor can actually approve.
- A hardening pass on the worst stories, sized and scheduled as real work rather than absorbed into refinement. Run it as a workshop with the product owner and the team; the rewrites are the coaching.
- A re-baselined date built on scored stories, so the new date rests on different evidence than the one that slipped.
- A re-score cadence, so movement is visible every fortnight rather than asserted at the end.
How do I keep everyone’s trust while doing this?
Make the findings an artefact rather than a conversation. A written report is re-readable, forwardable to a steering committee that wasn’t in the room, and impossible to remember as more hostile than it was. It also survives the staff change that happens on most troubled projects.
![A report cover page under a firm's own logo, titled “Backlog Quality Review”, with the line “Prepared by [Your agency]”. The headline reads: Hand the client a report with your name on it.](/_next/image?url=https%3A%2F%2Fcdn.vindexgateway.com%2FAgencyCarousel-6-report.png&w=3840&q=75)
Then close on the re-score, because a rescue is judged on movement rather than on diagnosis. When the second round shows the remaining backlog above the bar and the new date holding, the assessment stops being the meeting where things were bad and becomes the point where the project turned. Vindex adds that round-over-round comparison to the report automatically once a project has two scoring rounds; how a project works shows the loop end to end, and auditing a client backlog covers the cleaner start-of-engagement version of this exercise.
Related guides
The white-labeled backlog assessment report, section by section, with the real sample PDF as the reference: what each section shows, why it is there in a coaching or training engagement, what you can change, and what the client receives.
The baseline assessment a consultancy, coach, or trainer runs at the start of an engagement: what to ask for, why to score every story instead of sampling, how to package the findings as a deliverable, and how to re-score to show the change.
How a coach or consultant delivers hard findings about a client's backlog with the product owner in the room: rest on a stated standard, lead with aggregates, anonymize authors, pair every problem with a rewrite, and close on the re-score.
