What does a low-scoring story actually cost?
6 min read · Updated
A low-scoring story costs whatever the defect it becomes costs, paid later and by someone who never saw the story. That is the whole argument for scoring a client’s backlog before you coach from it, and it is easier to show than to explain, so here is one story from vague to shipped defect.
This is the guide to hand a product owner who agrees their requirements are “a bit thin” but doesn’t see why that should change how the team works, and the pre-read to send before a definition-of-ready workshop.
What does a story that looks ready look like?
Exactly like a story that is ready. That’s the problem: there is no visual difference between a well-specified story and an underspecified one, because what’s missing is absent, and absence doesn’t catch the eye in a backlog of four hundred rows. It is also why “we’ve always worked this way” is sincere: the team has never seen the gap, only its consequences, and the consequences arrive too late to be connected back.

Read it again as the developer who picks it up. What counts as a preference? Which channels? Does an admin override a user’s choice? What happens to preferences when the account changes? The story doesn’t say, so every one of those becomes a decision made quietly, at speed, by whoever is closest to the keyboard.
What does scoring actually flag?
A score turns that absence into something visible and sortable. The same story, scored, stops being a row in a backlog and becomes a specific, named risk with reasons attached, and the reasons come from a neutral engine rather than from the coach in the room.

The three findings matter more than the number. “Acceptance criteria are missing” is a fact anyone can verify in two seconds, which means the finding survives an argument. A single aggregate score invites debate about methodology; a finding that names the missing field invites a fix. For a coach this is the difference between an opinion the product owner can wave off and an observation they can check themselves.
What happens if nobody acts on it?
The story gets built. It has to: it was in the sprint, it looked ready, and no gate stopped it. The undefined behaviour becomes defined by whatever the implementation happens to do, and the first person to discover the mismatch is a user.

Notice what the defect is not. It isn’t a typo or a failed test. It is a behaviour nobody specified, discovered in production, now costing a triage cycle, a fix, a regression pass, a release, and a conversation with a stakeholder about how it shipped. All of it traceable to an empty field two sprints earlier, and none of it visible in the velocity chart the team uses to judge itself.
Why score requirements instead of tracking defects?
Because every other quality signal a team or its coach has arrives too late to change anything. Defect counts, escape rates, and rework hours are all measured after the build, when the money is already spent and the retrospective is about blame. Requirement quality is measurable at write time, which is the one moment when the fix is a paragraph and the coaching conversation is about a story rather than a failure.
The economics are the entire point. Adding acceptance criteria to PLT-2841 during refinement is a five-minute conversation. Fixing PLT-3190 after release is a defect ticket, a code change, a regression cycle, and an apology. Same missing information, two very different bills, and only one of them shows up in a sprint review as “unplanned work.”
Who ends up paying?
The team, in rework it never planned; the sponsor, in a date that moves; and you, in credibility, because you said the stories were thin and had nothing to point at when the team said they were fine. The party who writes the vague story is rarely the party who absorbs its cost, and that gap is exactly where an outsider’s warning gets lost. A score closes the gap: it attaches the cost to the story, at the moment the story is written, in a form the author can act on.
Which is also why the measurement has to happen before the coaching starts. Scoring the backlog at the baseline converts “the stories feel weak” into a named, counted risk; auditing a client backlog covers how to run that baseline, and for firms that also take fixed-fee delivery work, protecting fixed-bid margins covers the same chain from the vendor’s side.
How do I show a team this chain in their own backlog?
Score it and point at the worst three. The argument lands differently when the story on screen is one their team wrote, about their product, with their ticket key. Drop a CSV export into the free evaluator for real scores in about 30 seconds with no signup, then read presenting findings to a client before you put any of it in front of the team that wrote it. In a workshop, the three stories and their rewrites are the exercise.
Related guides
The white-labeled backlog assessment report, section by section, with the real sample PDF as the reference: what each section shows, why it is there in a coaching or training engagement, what you can change, and what the client receives.
The baseline assessment a consultancy, coach, or trainer runs at the start of an engagement: what to ask for, why to score every story instead of sampling, how to package the findings as a deliverable, and how to re-score to show the change.
How a coach or consultant delivers hard findings about a client's backlog with the product owner in the room: rest on a stated standard, lead with aggregates, anonymize authors, pair every problem with a rewrite, and close on the re-score.
