Measuring the Wrong Things: How Sprint Metrics Manufacture Progress Without Delivering It
Photo: Dr ian mitchell, CC BY-SA 3.0, via Wikimedia Commons
There is a scene familiar to anyone who has spent time in a software organization: the end-of-sprint review. Velocity is up. Story points are trending in the right direction. The burndown chart looks clean. Leadership nods approvingly. Engineers return to their desks. And somewhere in the codebase, a critical feature remains half-finished, a performance regression goes undetected, and a user-facing bug enters its third week without a fix.
This is not a hypothetical. It is a pattern that plays out in engineering organizations across the country, week after week, in companies that have invested heavily in agile tooling, sprint ceremonies, and productivity dashboards. The metrics look good. The product is not improving. The disconnect between the two is not accidental — it is structural.
The Appeal of Quantification
The impulse to measure engineering productivity is entirely understandable. Software development is expensive, timelines are uncertain, and stakeholders reasonably want assurance that engineering investment is generating returns. When a metric exists — when a number can be placed in a slide deck or a dashboard — it creates the appearance of rigor. It suggests that someone is watching, that progress is being tracked, that accountability exists.
The problem is that the most commonly used engineering metrics are measuring the wrong things. Story points quantify estimated effort, not delivered value. Velocity measures how many story points a team completes per sprint, which is to say it measures how many estimates a team completes — a recursive exercise that says very little about whether anything meaningful was built. Commit frequency tells you how often code was written, not whether the code was correct, maintainable, or necessary.
These metrics are not useless in every context. Within a stable team working on well-understood problems, velocity can be a reasonable planning tool. But when they are used as proxies for productivity, quality, or business impact, they create incentives that actively undermine the outcomes they purport to measure.
How Optimization Corrupts the Signal
When engineers are evaluated on metrics, they optimize for those metrics. This is not a character flaw — it is a rational response to incentive structures. When velocity becomes a performance indicator, story point estimates inflate. When commit frequency is tracked, engineers break work into smaller chunks. When sprint completion rates are reported to leadership, teams select work that is easy to complete rather than work that is important to complete.
This optimization process is sometimes called Goodhart's Law, after the British economist who observed that when a measure becomes a target, it ceases to be a good measure. In software engineering, the phenomenon is ubiquitous. Teams that are being measured on velocity learn to negotiate point estimates upward during planning. Teams under pressure to hit sprint commitments defer complex, high-value work in favor of smaller tickets that can be closed quickly.
The result is what might be called velocity theater: a performance of productivity that satisfies the metrics while the actual state of the product stagnates or deteriorates. The numbers continue to rise. The software does not.
What Genuine Progress Looks Like
If story points and velocity are insufficient, what should engineering teams measure instead? The answer depends on what the team is actually trying to accomplish, but several indicators consistently correlate with genuine engineering productivity and software quality.
Cycle time — the elapsed time between when work begins and when it is delivered to users — is a more honest measure of team efficiency than velocity. It accounts for the full complexity of delivering software, including review, testing, integration, and deployment, rather than just the estimation and development phases.
Deployment frequency and change failure rate, two of the four key metrics identified in the DORA research conducted by Google Cloud, offer insight into both the team's ability to ship and the reliability of what they ship. A team that deploys frequently with a low failure rate is demonstrably more capable than a team with high velocity and frequent rollbacks.
Mean time to recovery measures how quickly a team can restore service after an incident. This metric reflects the quality of the team's operational practices, monitoring infrastructure, and incident response processes — none of which are visible in a velocity chart.
Code quality indicators — including test coverage trends, static analysis findings, and the rate at which new defects are introduced — provide a window into the long-term health of the codebase that story points fundamentally cannot.
The Organizational Courage Required
Shifting from vanity metrics to meaningful ones requires something that is genuinely difficult in most organizations: the willingness to accept uncertainty and ambiguity in exchange for accuracy. Story points and velocity feel precise. Cycle time and deployment frequency feel less controllable. Leadership accustomed to clean dashboards may resist measurements that reveal more complexity than they resolve.
This resistance is worth confronting directly. The purpose of engineering metrics is not to make leadership comfortable — it is to help engineering teams understand what is working and what is not, and to give the organization an accurate picture of where value is being created. Metrics that obscure this picture are not merely neutral; they are actively harmful, because they allow organizational dysfunction to persist behind a facade of apparent productivity.
Engineering leaders who want to make this shift should begin by having an honest conversation with their teams about what the current metrics are actually measuring and what they are missing. This conversation is often revelatory. Engineers typically know, with considerable precision, which metrics are being gamed and how. They are frequently relieved to be asked.
Toward a More Honest Accounting
None of this is an argument against measurement. Engineering teams that operate without feedback loops are flying blind, and good metrics genuinely improve decision-making. The argument is for measurement that is honest — that reflects the actual complexity of building software rather than reducing it to numbers that are easy to track and easy to manipulate.
The teams that build the best software are not the ones with the highest velocity. They are the ones with the clearest understanding of what they are building, why they are building it, and how well it is working once it is in users' hands. The metrics that support that understanding are available. The decision to use them — and to stop performing productivity for its own sake — is one that engineering leadership must make deliberately.
Sprint metrics will continue to lie to you for as long as you allow them to. The question is whether the performance is worth the cost.