Your DORA Metrics Look Great. Your Delivery Risk Might Be Getting Worse.
A VP of Engineering at a 140-person SaaS company pulls up the quarterly board deck. Deployment frequency is up 60% year over year. Lead time for changes has dropped from four days to under six hours. Change failure rate sits at 9%, comfortably inside the range DORA's research associates with elite performers. Mean time to restore is measured in minutes, not hours. Every number on the slide is trending the right direction, and has been for three straight quarters.
Two weeks after that board meeting, a routine failover test reveals that the last four "restores" logged under two minutes were not restores at all. They were feature flags flipped back off before the change had finished rolling out to more than 3% of traffic — incidents that never touched most customers, closed before they ever generated a page, and therefore never logged as incidents in the tracking system the dashboard pulls from. The actual mean time to restore, measured against customer-visible incidents the support team can point to, is closer to ninety minutes. Nobody lied in the board meeting. Nobody fabricated a number. The dashboard was accurate, by its own definitions, the entire time. It was also almost useless as a description of the engineering organization's actual risk profile.
This is not a story about dishonest engineers. It is the ordinary, statistically expected outcome of what happens when a well-designed measurement system gets attached to consequences its designers never intended it to carry. DORA metrics — deployment frequency, lead time for changes, change failure rate, and time to restore service — were built by researchers to help engineering organizations understand their own delivery capability over time. They were never designed to be compared across teams, tied to individual performance ratings, or presented as a scorecard for a board that will draw funding, hiring, or leadership conclusions from them. The moment they get used that way, the incentive structure around them changes completely, and the numbers start to drift away from the reality they were built to represent.
This article is not an introduction to DORA metrics. If you are reading it, you almost certainly already report deployment frequency, lead time, change failure rate, and MTTR upward through some kind of dashboard, quarterly review, or board deck. The argument here is narrower and more uncomfortable: those numbers are very likely being gamed right now, in your organization, by people who are not acting in bad faith, and the dashboard has no mechanism to tell you this is happening. What follows is an anatomy of how that gaming actually occurs metric by metric, three labeled hypothetical scenarios that show the mechanism in enough detail to recognize it, and a structural framework — not a stricter policy, not a sterner memo — for telling whether your own numbers are still measuring what you think they measure.
The Measurement Paradox Behind Every Healthy-Looking Dashboard
The mechanism at work here has a name, and it predates software engineering by decades. In 1975, British economist Charles Goodhart wrote a paper on UK monetary policy observing that "any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes." Anthropologist Marilyn Strathern later condensed the idea into the phrasing now attached to Goodhart's name: "When a measure becomes a target, it ceases to be a good measure." The insight is not that people cheat. It is that a statistical relationship between a proxy (deployments per week) and a real goal (fast, safe delivery of value to customers) only holds under the conditions in which nobody is deliberately steering toward the proxy. Once the proxy becomes the thing that gets rewarded, watched, or compared, people — quite rationally — start optimizing for the proxy instead of the goal, and the correlation between the two breaks down.
A parallel and slightly earlier formulation comes from psychologist Donald T. Campbell, who wrote in a 1979 paper on evaluating social programs: "The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor." Campbell was writing about standardized test scores in education, not software delivery pipelines, but the mechanism transfers almost without modification. He illustrated it with achievement tests: valuable as a general indicator of school performance under normal teaching conditions, but "when test scores become the goal of the teaching process, they both lose their value as indicators of educational status and distort the educational process in undesirable ways." Replace "test scores" with "deployment count" and "teaching process" with "release process," and the sentence still parses cleanly for a modern engineering organization.
Neither Goodhart nor Campbell was describing a moral failure. Neither was warning that people become dishonest when they are measured. The warning is structural: any proxy metric, no matter how well designed, degrades in accuracy the moment optimizing that specific number carries consequences separate from the underlying behavior it was meant to represent. DORA's own research team has published explicit guidance anticipating exactly this failure. Their public documentation on the four keys metrics states plainly that organizations should avoid treating the metrics as fixed targets, because doing so "invokes Goodhart's Law and encourages manipulation rather than genuine improvement," and that the metrics are meant to help a team or service improve relative to its own baseline — not to be compared across teams, ranked competitively, or used to evaluate individuals. Most organizations that report DORA metrics upward violate at least one of those explicit guardrails, usually without realizing the guardrail exists.
The distortion is rarely conscious sabotage. It is the sum of dozens of small, locally reasonable decisions made under pressure by people who are not trying to deceive anyone. An engineer who splits one meaningful change into three commits before a sprint review is not lying — she is responding, correctly, to an environment that rewards a higher commit count. A team lead who reclassifies a customer-impacting outage as a "planned mitigation" rather than an incident is not falsifying records in the criminal sense — he is applying a definition that the organization itself left ambiguous, in the direction that is least uncomfortable for his team. Multiply either behavior across a 200-engineer organization over four quarters, and the aggregate dashboard drifts substantially from the underlying reality, even though no single decision along the way looks like fraud from the inside.
This matters because the standard organizational response to a metric that looks wrong is to tighten the definition, add an audit step, or threaten consequences for miscategorization. All three responses treat the symptom as an integrity problem to be policed rather than a measurement-system design flaw to be corrected. Tightening a gamed metric without changing what it is used for produces, at best, a temporary improvement followed by a more sophisticated form of the same gaming. The rest of this article works through why that is true, what the gaming actually looks like at the level of individual engineering decisions, and what a different design for the measurement system looks like in practice.
Anatomy of a Gamed Metric: How Each DORA Indicator Gets Distorted
Each of the four core metrics has a distinct failure mode, driven by a distinct incentive. Understanding the mechanism for each one is what makes it possible to recognize the pattern in your own organization's numbers rather than treating a healthy dashboard as reassurance.
Deployment Frequency: Counting Motion Instead of Progress
Deployment frequency was designed to measure an organization's capacity to release small batches of change safely and often — a proxy for reduced batch-size risk, tighter feedback loops, and lower blast radius per release. It becomes a target the instant leadership starts citing "deployments per day" in a board deck or ties a team's release cadence to a quarterly OKR.
Once that happens, three adaptations appear with predictable regularity. First, teams redefine what counts as a deployment. A configuration flag toggle, a documentation update to a service repository, or a no-op canary release with zero customer-facing change all start getting logged as deployments, because the tracking tool counts merges to a deploy branch or pipeline runs rather than customer-facing releases. Second, engineers split changes that would previously have shipped as one coherent unit into several smaller commits or pull requests, each triggering its own pipeline run, purely to inflate the count — not because smaller batch sizes improve safety in this instance, but because the counting mechanism rewards the split regardless of whether it improves anything. Third, and most consequential, teams start deploying on a schedule dictated by the metric rather than by readiness. A team under pressure to maintain a "daily deploy" streak will sometimes ship a change that is not fully validated, simply to avoid a gap in the graph — inverting the entire purpose of the metric, which was originally a proxy for confidence and safety, not obligation.
The tell is a dashboard where deployment count keeps climbing while the average size of each deployment (lines changed, files touched, or scope of the associated ticket) keeps shrinking toward near-zero, or where a disproportionate share of "deployments" touch configuration and feature flags rather than application code.
A new wrinkle: AI-assisted development and inflated activity counts. AI coding assistants have made it materially easier to produce a high volume of small, plausible-looking commits and pull requests with limited additional engineering effort, which changes the economics of deployment-frequency gaming without changing the underlying mechanism. An engineer working with an AI pair-programming tool can generate several independent, narrowly scoped pull requests — a refactor here, a small utility function there, an expanded test file — each one legitimate in isolation, in roughly the time it used to take to produce one. None of this is dishonest on its face; AI-assisted development genuinely can support smaller, safer batches when used deliberately. The risk is that it lowers the cost of producing countable activity to nearly zero, which means a deployment-frequency target that used to require real organizational strain to game now requires almost none. A team chasing a cadence number can generate a defensible-looking stream of AI-assisted micro-commits without materially advancing the roadmap, and the pattern is harder to spot than the manual version, because each individual pull request looks like ordinary, well-scoped engineering work rather than an obvious cosmetic split. The diff-size counter-metric introduced later in this article becomes more important, not less, in an environment where the marginal cost of producing an additional countable commit keeps falling.
Lead Time for Changes: Moving the Starting Line
Lead time for changes measures the interval from code committed to code running in production — a proxy for how much friction exists in your build, review, test, and release pipeline. Under pressure, the most common distortion is not speeding up the pipeline. It is moving when the clock starts.
Work that used to be captured inside the measured interval — design discussion, spec review, early prototyping, exploratory branches — gets pushed into an unmeasured "pre-commit" phase that can stretch for weeks without ever appearing on the lead-time chart. Engineers hold a branch locally, iterate through several rounds of self-review, and only open the pull request once the change is essentially final, at which point it sails through review and CI in under an hour. The dashboard reports outstanding lead time. The customer, the roadmap, and the sprint plan all experienced the true, much longer cycle time, because the actual constraint — early ambiguity, unclear requirements, slow design sign-off — never got measured or improved. A second, subtler version of the same distortion happens with large, long-running feature branches: work sits unmerged for six weeks, then gets broken into a burst of small commits merged within an hour of each other right before a reporting period closes, producing an artificially short lead-time average for that reporting window while doing nothing to address the actual six-week delay the customer experienced.
The tell is a lead-time chart that looks excellent while cycle time — the interval from when work is picked up to when it ships, measured independently through project-tracking timestamps rather than commit timestamps — tells a very different story.
Change Failure Rate: Redefining "Failure" Until It Disappears
Change failure rate — the share of deployments that require a hotfix, rollback, or immediate remediation — is the metric most directly exposed to definitional gaming, because "requires immediate intervention" is a judgment call, and the people making that judgment are frequently the same people whose performance the judgment reflects.
The gaming here rarely takes the form of an outright false statement. It takes the form of quietly narrowing what counts as a qualifying failure. A partial outage affecting 4% of traffic for eleven minutes gets classified as "degraded performance" rather than an incident because no formal severity level was triggered by the on-call runbook. A rollback initiated within the same deploy window, before the change technically reached "fully released" status in the deployment tool, gets excluded from the failure count because the tool's data model treats it as an aborted deployment rather than a failed one. A hotfix shipped the same day gets logged as a "follow-up enhancement" rather than a remediation, because the ticket was filed by the same engineer who wrote the original change and there was no formal incident review to force a different classification. None of these are lies in isolation. In aggregate, across a year of ambiguous edge cases resolved consistently in the direction that keeps the number low, the reported change failure rate can end up materially disconnected from the rate of customer-visible problems following a release.
The tell is a change failure rate that stays flat or improves while customer support ticket volume tagged to recent releases, or the number of hotfix deploys within 48 hours of a preceding release, moves in the opposite direction.
Time to Restore Service: Stopping the Clock Early
Time to restore service measures how quickly a team recovers from a failure once it has begun. The most common distortion is declaring recovery the moment a symptom is suppressed rather than the moment the underlying cause is resolved and customer impact has genuinely ended.
A feature flag gets flipped off, a bad pod gets killed and replaced, or traffic gets rerouted to a healthy region — all valid immediate mitigations — and the incident gets marked resolved at that instant, because that is when the alerting stopped firing. The root cause, still present in the codebase, resurfaces two weeks later as what looks like an unrelated new incident, resetting the MTTR clock again rather than being counted as a continuation of the same underlying failure. A second common pattern: incidents that get caught and mitigated before they trigger the organization's formal severity threshold — say, a partial degradation affecting a small percentage of a specific customer segment — never enter the incident tracker at all, so they contribute zero minutes to the MTTR calculation despite having been a real, customer-visible failure that a real engineer spent real time diagnosing and fixing.
The tell is an MTTR trend that improves steadily while the count of distinct incidents mysteriously narrows at the same time low-severity, sub-threshold tickets from support or customer success keep climbing — suggesting failures are being caught earlier in visibility but not necessarily resolved faster in substance, or are being kept below the threshold that would require them to be logged at all.
Why Capable, Well-Intentioned Engineers Do This
It is tempting to treat each of the four patterns above as an integrity failure requiring a stricter policy or a values conversation. That diagnosis is usually wrong, and acting on it usually makes the problem worse, because it treats a structural incentive problem as a character problem.
Consider the position of an engineering manager whose team's deployment frequency and change failure rate appear, unlabeled, on a spreadsheet the VP of Engineering reviews before calibration season. The manager did not design the incentive. She did not choose to have her team's numbers compared against three other teams working on systems of very different age, complexity, and customer criticality. She is, however, the person who will sit across from her own reports in a few weeks and decide who gets a strong rating, and she knows the aggregate team number will color how that conversation is received two levels up. Given that reality, narrowing the definition of "failure" by a few edge cases, or encouraging her team to ship in slightly smaller increments to keep the deployment count healthy, is not a moral lapse. It is a locally rational response to a system that has made the proxy consequential without making the proxy accurate.
This is precisely the pattern both Goodhart and Campbell were describing, and it is why punitive responses tend to fail. If leadership discovers gaming and responds by threatening consequences for "cheating" the metrics, the rational adaptation is not to stop optimizing the proxy — it is to get better at making the optimization invisible. Definitional gaming that used to happen in the open, discussed candidly in a retro, moves underground. The dashboard keeps looking healthy. The gap between the dashboard and reality gets larger, not smaller, because the people closest to the truth now have a reason to hide it rather than surface it.
The more productive framing, and the one this article works from, is that gaming is evidence the measurement system's design is broken, not evidence that the people inside it are dishonest. A system that ties comparative, individual, or high-stakes consequences to a proxy metric will produce proxy optimization from rational actors regardless of their character. The fix has to change the system, not the actors.
The pattern is not unique to software, and it is worth naming the most consequential real-world example outside engineering to show how far the same dynamic can travel once a proxy metric acquires real stakes. Wells Fargo's retail bank spent years setting aggressive cross-sell targets for the number of accounts and products each employee opened per customer, tied directly to compensation and continued employment. Employees, under sustained pressure to hit a number that leadership treated as a proxy for genuine customer relationship depth, opened millions of accounts customers had never requested — a scandal that became public in 2016 and led to billions of dollars in regulatory penalties and a congressional hearing. Almost nobody involved at the branch level set out to defraud customers on day one. The incentive structure made the proxy the thing that determined whether they kept their job, and the proxy stopped correlating with the underlying goal it was built to represent. Engineering organizations rarely produce headlines on that scale, but the mechanism connecting a target, a consequence, and a slow drift away from the truth the target was meant to track is identical. A change failure rate quietly redefined downward and a cross-sell count quietly inflated are the same failure wearing different clothes.
The Statistical Trap Hiding Inside Small Teams' Numbers
A separate and less discussed problem compounds the gaming risk: most engineering teams are far too small for their DORA metrics to be statistically stable from one reporting period to the next, which makes ordinary noise easy to mistake for a trend, and easy to present as one.
A team of eight engineers shipping fifteen deployments in a typical month, with two "failed" deployments requiring a hotfix, reports a 13% change failure rate. The following month, with the same underlying engineering quality and the same underlying risk, one additional deployment happens to require a rollback purely by chance — a dependency version conflict nobody could reasonably have caught, unrelated to process quality — and the change failure rate jumps to 20%. Nothing about the team's actual capability changed. The number moved by more than half, because the denominator was never large enough to produce a stable percentage in the first place. A leadership team unaware of this will often respond to the jump with a review, a process change, or a pointed question in a one-on-one, all aimed at a signal that was mostly statistical noise.
This dynamic creates a second, quieter incentive to game the numbers that has nothing to do with deliberate manipulation: once a team has been burned by a leadership reaction to a noisy monthly swing, the rational response is to smooth the reported number artificially — delaying a hotfix log entry into the following period, or absorbing a marginal incident into an existing one — simply to avoid triggering another review over what the team correctly understands to be statistical variance rather than a real signal. The fix is not more scrutiny of small swings. It is reporting these metrics over rolling quarterly windows rather than monthly snapshots for any team below roughly twenty deployments per period, and treating a single period's movement as a question to investigate rather than a conclusion to act on.
A concrete worked calculation makes the instability easier to see. Suppose a team ships 18 deployments in a month, two of which require a hotfix — a change failure rate of 11.1%, comfortably inside "elite" range. The following month, the team ships the same 18 deployments, and one additional borderline event — a partial degradation that a different on-call engineer would have classified differently — gets counted as a third failure instead of being absorbed into an existing incident. The rate jumps to 16.7%, a swing of more than five and a half points driven by exactly one classification decision on one edge case. Reported as a single number in a board deck, that swing reads as a meaningful deterioration in release quality. Reported alongside its sample size and the specific edge-case decision that drove it, it reads as exactly what it is: normal variance at low volume, amplified by an ambiguous classification call. The percentage format itself, precise to one decimal place, disguises how few actual events are driving the number.
The Compliance and Audit-Trail Blind Spot
Gamed delivery metrics carry a consequence beyond misleading a board deck: for organizations subject to SOC 2, ISO 27001, PCI DSS, or contractual customer SLAs, the same incident and deployment records feeding the internal dashboard often double as the evidence base an external auditor or a customer's security team will eventually request. A change-management control that requires evidence of testing and rollback procedures for production changes typically gets satisfied by pointing to the deployment and incident tracking system — the same system whose definitions this article has described drifting under reporting pressure.
This creates a specific and underappreciated risk. An auditor sampling deployment records for evidence of a functioning change-management process will generally accept whatever the organization's own system classifies as a deployment or an incident, because auditors test whether a documented control is being followed consistently, not whether the underlying classification itself is honest. A change-failure-rate definition that has been narrowed to exclude same-window rollbacks does not just distort a leadership metric — it can produce an audit trail that understates the frequency of production incidents to the very party responsible for independently verifying the organization's control environment. This is rarely intentional audit evasion; it is the same locally rational classification behavior described throughout this article, applied to a system that happens to also serve a compliance function. The practical implication is that the definition-ownership fix recommended later in this article is not only a metrics-integrity improvement. For regulated or audited organizations, it is also a control-environment integrity issue, and worth raising directly with whoever owns the compliance relationship, not just with engineering leadership.
Case Scenario One: The Cadence Tax
The following is a hypothetical, illustrative scenario. It does not describe a real QAtronic client, and all figures are constructed for explanatory purposes only.
Picture a 60-engineer SaaS company selling workflow software to mid-market logistics firms. Following a board recommendation, engineering leadership adopts DORA metrics as the primary delivery health indicator, and sets a target: every product team should deploy to production at least once per business day. The stated goal is legitimate — the company's release process had genuinely been slow, with some teams shipping only every two to three weeks, and smaller, more frequent releases would reduce risk per deploy if the shift happened for real reasons.
The hidden assumption behind the target is that deployment cadence and batch size are the binding constraint on delivery risk for every team, regardless of what that team is building. In practice, one team owns the billing and invoicing subsystem, where changes routinely require coordination with the finance-in-tax integration, a slower external vendor, and cannot safely ship in small independent increments without risking half-applied billing logic reaching production. Rather than raise this with leadership — a conversation that would read, in a target-driven environment, as an excuse — the team lead instructs engineers to break billing changes into cosmetic-looking commits (renaming variables, adjusting comments, minor refactors with no functional change) interleaved between the substantive commits, each one triggering the CI/CD pipeline and counting as a deployment. Within two months, the billing team's deployment frequency matches the rest of the organization.
The consequence surfaces four months later. A genuinely substantive billing change — updated tax-calculation logic required by an upcoming regulatory deadline — gets shipped in the same pattern of small, frequent commits the team has been using to hit cadence, except this time the fragmentation obscures a subtle ordering dependency between two of the commits. The dependency would have been obvious in a single coherent code review of the full change; split across five smaller reviews spread over three days, no single reviewer saw the whole picture. The bug reaches production, miscalculates tax on a subset of invoices for eleven days before a customer flags a discrepancy, and the finance team spends three weeks manually reconciling affected accounts.
The decision facing leadership after the postmortem is not "how do we punish the billing team for gaming the metric." It is whether deployment cadence should have been a uniform target across systems with fundamentally different risk profiles in the first place. The better approach, adopted after this scenario, sets cadence expectations per system based on that system's actual coupling and blast radius, drops the cross-team comparison entirely, and adds a simple structural safeguard: a pull request's diff is compared against its linked ticket scope, and PRs with a suspiciously high ratio of file touches to logical lines changed get flagged for review — not as a punishment mechanism, but as a signal that the metric and the work have started to diverge.
Building a Detection Framework: Pairing Every Metric With a Counter-Metric
The single most useful structural change available to a leadership team is also the least dramatic: stop reading any DORA metric in isolation, and pair each one with a counter-metric drawn from a different data source that would be difficult to game through the same mechanism, because it is generated by a different system, a different team, or a different set of incentives.
This idea is not unique to this article — it echoes the SPACE framework for developer productivity published by researchers including Nicole Forsgren, which explicitly warns that activity metrics such as commit or deployment counts are "the most ubiquitous productivity measure and often the most misused" precisely because they can be inflated without reflecting genuine progress, and recommends drawing on at least three different dimensions of measurement before drawing conclusions. The counter-metric framework below applies that same triangulation logic specifically to the four DORA indicators.
| Reported metric | Common gaming mechanism | Independent counter-metric to pair it with | What divergence signals |
|---|---|---|---|
| Deployment frequency | Splitting trivial or cosmetic changes into separate deploys; counting config/flag toggles as deployments | Median lines of code (or logical diff size) per deployment, tracked over time | Rising deploy count with falling median diff size suggests fragmentation for cadence, not genuine batch-size improvement |
| Lead time for changes | Delaying the pull-request open date until work is functionally complete, hiding pre-commit time | Cycle time from ticket "in progress" status to production, measured in the project tracker independent of git timestamps | A large, growing gap between lead time and cycle time suggests work is happening off the measured clock |
| Change failure rate | Narrowing the definition of "failure" to exclude sub-threshold incidents, same-window rollbacks, or quick hotfixes | Customer support tickets referencing a recent release, tagged independently by the support team | Flat or improving change failure rate alongside rising release-linked support volume suggests failures are being reclassified, not reduced |
| Time to restore service | Marking an incident resolved at first symptom suppression rather than root-cause resolution | Recurrence rate: incidents reopened or repeated within 30 days on the same underlying component | Improving MTTR alongside a rising 30-day recurrence rate suggests symptoms are being treated, not causes |
None of the four counter-metrics in the right-hand columns is exotic or expensive to collect. Diff size is available from any git host. Cycle time is available from any project tracker that records status-change timestamps. Release-linked support tickets require nothing more than a tagging convention between the support and engineering tooling. Recurrence rate requires only that incidents be linked to a component or root-cause tag rather than treated as unrelated events. The reason this pairing works as a detection mechanism is structural, not statistical: each counter-metric is generated by a system, and often a team, with a different incentive structure than the one generating the primary metric, which makes it much harder for the same local optimization to distort both numbers in the same direction at the same time.
Case Scenario Two: The Incident That Wasn't
The following is a hypothetical, illustrative scenario, constructed for explanatory purposes. It does not describe a real QAtronic client.
Consider an enterprise software vendor with roughly 300 engineers, serving large logistics and manufacturing customers on multi-year contracts with contractual uptime commitments. Change failure rate has been the most closely watched metric on the engineering leadership dashboard for two years, because it feeds directly into a customer-facing reliability report sent quarterly to the company's largest accounts. Engineering leadership has, with good intentions, made it clear that a rising change failure rate would trigger a formal review of the release process for the team responsible.
The hidden assumption is that the incident-classification process, run independently by each team's on-call engineer at the moment of the event, will apply a consistent standard regardless of which team is involved or what a "failure" would mean for that team's standing. In practice, the formal incident-severity rubric requires a customer-facing impact affecting more than 5% of active sessions for at least ten minutes before an event is logged as a qualifying incident in the system that feeds the change failure rate calculation. A partial outage affecting 3% of sessions for eight minutes, caused directly by a recent deployment, technically falls below that threshold. Over several quarters, on-call engineers — under no explicit instruction to hide anything, simply applying the rubric as written, in the direction that avoids triggering an uncomfortable review for their own team — begin resolving borderline events quickly and quietly rather than formally declaring them, because a formally declared incident starts a clock and a paper trail that a quietly resolved one does not.
The consequence surfaces when one of the largest customers, whose account team tracks its own uptime experience independently for contractual purposes, reports six distinct service disruptions over a quarter in which the internal change failure rate dashboard shows zero qualifying incidents for that customer's service tier. The customer's data and the internal dashboard are both technically accurate — they are simply measuring different thresholds, one set by contract, one set by an internal rubric that had quietly become a target rather than a genuine classification standard.
The decision facing leadership is not whether to discipline the on-call engineers who made a defensible judgment call under an ambiguous rubric applied consistently over time. It is whether a self-reported classification system, in which the team whose performance is measured by the classification is also the team making the classification, can ever be trusted once the classification carries consequences. The better approach removes on-call engineers from the sole authority over incident classification for anything that will feed a customer-facing or leadership-facing metric, and instead sets the qualifying threshold independently by pairing internal incident logs against customer-reported disruptions and automated availability monitoring — sources the on-call engineer does not control and has no reason to shade in either direction.
Case Scenario Three: The Elite Performer That Wasn't
The following is a hypothetical, illustrative scenario, constructed for explanatory purposes. It does not describe a real QAtronic client.
Picture a fintech company processing payment reconciliation for small business customers, roughly 90 engineers, growing quickly, with a board that has started asking pointed questions about engineering velocity relative to headcount growth. The CTO, wanting a clean, defensible answer, adopts the DORA benchmark categories directly and sets a goal: reach "elite" classification (multiple deploys per day, sub-hour lead time, change failure rate under 15%, MTTR under an hour) within two quarters, with progress reviewed at each engineering leadership sync and summarized for the board each quarter.
The hidden assumption is that hitting the elite thresholds on all four metrics simultaneously, on this specific timeline, is achievable through genuine engineering improvement — faster CI, smaller batches, better test automation, more reliable rollback tooling — rather than through selective redefinition under time pressure. Real improvement does happen on some dimensions: the CI pipeline genuinely does get faster after a focused investment in parallelized test execution. But the aggressive timeline creates pressure that the genuine improvements alone cannot satisfy on schedule, and two separate distortions compound each other in the same reporting period. First, several teams begin closing incidents as soon as a mitigating flag flip suppresses customer-visible symptoms, moving the MTTR number down sharply, consistent with the "stopping the clock early" pattern described earlier. Second, and more consequentially, the definition of "deployment" in the tracking dashboard is quietly updated — by a well-meaning platform engineer trying to make the tooling more accurate, not to game anything — to exclude infrastructure and database-migration changes from the deployment count, on the reasoning that these are "operational, not product deployments." This change alone reclassifies several previously "failed" deployments (migration rollbacks) out of the change-failure-rate denominator entirely, because failed migrations were no longer counted as deployments at all. The change failure rate improves by several percentage points in the same quarter, for reasons that have nothing to do with any actual improvement in release safety.
The consequence surfaces roughly five months later, when a payment reconciliation batch job — the kind of infrastructure-adjacent change that had been quietly excluded from the deployment count — ships with a currency-rounding defect that goes undetected for nine days because it no longer triggers the same monitoring or review rigor that "counted" deployments receive. By the time it is caught, several thousand small business accounts have received reconciliation reports with cumulative rounding discrepancies, requiring a multi-week manual correction process and formal customer notifications.
The decision facing the CTO in the aftermath is not simply whether to reverse the definitional change (which does need to happen). It is a harder question: whether tying an aggressive, board-visible timeline to a specific benchmark classification created the pressure that made both distortions — the MTTR shortcut and the quiet redefinition — more likely to occur, independent of anyone's intent to deceive. The better approach treats the DORA benchmark categories as a general orientation for identifying where the organization's delivery capability sits relative to published research, not as a compliance target with a board-visible deadline, and separates "what counts as a deployment" from any team empowered to benefit from narrowing that definition — that decision now sits with a cross-functional review that includes QA and reliability engineering, not the platform team alone.
Redesigning the Reporting System So the Incentive to Game It Disappears
Detection alone is not a fix. A leadership team that catches gaming, corrects the specific numbers, and changes nothing about how the metrics are used will see the same distortions reappear in a different form within a year, because the underlying incentive — a proxy metric with consequences attached — is unchanged. Four structural changes address the incentive directly rather than policing its symptoms.
Decouple the metrics from individual performance review entirely. DORA's own guidance is explicit that these metrics are intended to describe a team or service's delivery capability relative to its own history, not to evaluate individuals. The instant a metric appears, even implicitly, in a performance conversation — "your team's deployment frequency lagged this quarter" directed at a manager's own rating — every engineer on that team has a personal stake in the number moving in one direction, and the aggregate becomes far more likely to drift from reality than a number nobody's compensation touches. This does not mean delivery capability should be invisible to performance conversations. It means the connection should run through qualitative engineering judgment informed by the metrics, reviewed by someone with context on what the numbers actually represent for that specific system, rather than through the raw number itself appearing on a scorecard.
Stop comparing teams to each other on these metrics. A payments team maintaining a system with strict regulatory review requirements will structurally never match the deployment frequency of a marketing-site team shipping copy changes, and forcing them onto the same leaderboard creates pressure to distort rather than pressure to improve. DORA metrics were designed and validated for tracking a single team or service against its own baseline over time. Cross-team comparison was never part of the original research design, and using the metrics that way manufactures gaming pressure that would not otherwise exist.
Assign definition ownership to a party without a stake in the outcome. The single clearest pattern across all three case scenarios above is that the group whose performance a metric measures should not be the same group with sole authority to define what counts inside that metric. Whether that means a platform or QA function owns the canonical definition of "deployment," "incident," and "failure" across the organization, or an explicit change-control process requires cross-functional sign-off before any metric definition changes, the point is the same: definitional drift should require friction proportional to how much the definition affects a reported outcome, not zero friction because it happened to be one engineer's reasonable-sounding tooling decision.
Report ranges and context, not single numbers with false precision. A board slide that reports "change failure rate: 9%" invites a binary read — good or bad — that a single percentage cannot honestly support. A more honest version reports the number alongside its counter-metric, a note on what changed in measurement methodology since the last report, and a qualitative sentence from engineering leadership on what the number does and does not capture this quarter. This is a harder story to tell in a board meeting than a green arrow, and that difficulty is exactly the point: metrics that are hard to misrepresent are usually metrics that are honestly reported with their own limitations attached.
Put a sunset clause on any metric framed as a target rather than a trend. Targets accumulate gaming pressure the longer they stay in place, because every quarter a target survives, the population of people who have learned how to satisfy it without doing the underlying work grows. A "reach elite change failure rate within two quarters" goal, even if it starts as a genuine stretch aim, should either convert into an ordinary trend-line the organization simply keeps an eye on once achieved, or expire and get replaced with a different, harder target — not persist indefinitely as a standing scorecard item, which is exactly the condition under which Goodhart's dynamic compounds year over year. An organization that has been reporting the same target-framed metric for three or more consecutive years without ever revisiting whether it should still be a target at all is very likely reporting a number that has been optimized past the point of meaning anything.
Taken together, these five changes share a common thread: none of them require better engineers, more disciplined engineers, or a values campaign. They require leadership to change what the numbers are allowed to touch — compensation, comparison, deadline pressure, and single-party definitional authority — because those are the four levers that convert an honest proxy into a target worth gaming.
The Metric Integrity Audit: A Framework for Testing Your Own Dashboard
The following framework was built specifically for this article as a practical starting point for engineering leaders who want to test whether their own delivery dashboard has drifted from what it claims to measure. It is not a generic maturity model repurposed with new labels — it is a sequence of eight concrete checks, each answerable with organization-specific data most companies already have somewhere, even if it has never been assembled in one place before.
- Trace one metric back to its raw event source. Pick change failure rate. Pull the underlying list of deployments the calculation is based on for the last two quarters, and the list of events counted as "failures." Read the actual incident descriptions, not just the count. Ask whether a reasonable outside observer, with no stake in the number, would classify each borderline case the same way your team did.
- Compare deployment count against median diff size over the same window. If deployment frequency has risen while median lines-changed-per-deploy has fallen sharply, investigate whether the increase reflects genuine batch-size improvement or fragmentation for cadence.
- Compare git-based lead time against tracker-based cycle time for the same sample of tickets. Pull twenty recent tickets. Record when the ticket entered "in progress" in the project tracker and when the associated pull request was actually opened. A large, consistent gap indicates work is happening before the lead-time clock starts.
- Cross-reference the incident log against customer support tickets tagged to recent releases. If support-reported, release-linked issues are rising while the internal incident count is flat, treat this as the single highest-priority discrepancy to investigate, because it is the pattern most directly connected to real customer harm.
- Ask who has authority to define what counts. For each of the four metrics, identify the specific role or team that decides the boundary case — what counts as a deployment, an incident, a failure, a restoration. If that role is the same role whose performance the metric measures, treat this as a structural risk regardless of whether any specific number currently looks suspicious.
- Check whether the metric definitions have changed in the last four quarters, and who approved the change. A definitional change that quietly improved a reported number, approved without cross-functional review, is a stronger signal of drift than the number itself.
- Look for recurrence, not just resolution speed. Pull the list of incidents marked resolved in the last two quarters and check how many recurred, in a related form, on the same component within 30 days. A rising recurrence rate alongside an improving MTTR is close to definitive evidence that symptoms are being suppressed rather than causes fixed.
- Ask what would happen to any individual's standing if this specific number got worse next quarter. If the honest answer involves a specific person's compensation, rating, or standing, the incentive to distort that number already exists, independent of whether distortion has occurred yet. Decouple the consequence before assuming the number is currently accurate.
None of these eight checks requires new tooling. All eight can typically be completed by a VP of Engineering or a QA/reliability lead within a few days using data the organization already has scattered across its git host, project tracker, incident system, and support platform. The value is in the cross-referencing, not in any single data source — which is the same principle underlying the counter-metric pairing framework above.
Run the audit as a standing quarterly exercise rather than a one-time investigation, and rotate who performs it. A first audit typically surfaces the most obvious drift — a stale definition nobody has revisited in two years, an incident rubric one team interprets more loosely than the rest of the organization. Later audits, performed by someone who has learned where the previous round's blind spots were, tend to surface subtler and more organizationally revealing findings: which specific individuals have the most personal stake in a given number, and whether the informal social pressure around a metric has shifted since the last review. Treating the audit as a single cleanup event rather than a recurring discipline invites exactly the kind of definitional decay described throughout this article to reassert itself within a year or two.
Where This Framework Reaches Its Limits
Triangulation and structural decoupling are not universally proportionate responses, and applying heavy governance to every metric in every organization creates its own cost. A five-person engineering team at a pre-seed startup shipping to a handful of design-partner customers does not need a cross-functional metric-definition committee; the founder and the two other engineers already know, from direct daily contact, whether the last release caused a problem. Formal metric governance becomes valuable roughly at the point where the people setting delivery expectations are no longer the same people who can see the underlying work directly — typically somewhere in the 30-to-80-engineer range, though the actual trigger is organizational distance, not headcount alone. The table below summarizes how the risk and the proportionate response scale across company stages.
There is a second limit worth naming directly: none of this framework works if leadership is not genuinely willing to hear that a number they have already presented externally — to a board, an investor, or a customer — was wrong. The Metric Integrity Audit will occasionally surface a finding that is organizationally inconvenient, not just technically interesting: a change failure rate reported to the board last quarter as 9% that should honestly have been closer to 15% under a consistent classification standard. An organization that treats this discovery as a reason to quietly revert to the old definition rather than correct the record has not fixed the measurement problem; it has demonstrated, to everyone who ran the audit, that accurate numbers are less welcome than comfortable ones, which reintroduces exactly the pressure this article describes. The audit only produces lasting value in a culture willing to report a corrected, less flattering number upward and explain why the correction happened.
| Organization stage | How gaming typically shows up | Proportionate response |
|---|---|---|
| Early-stage startup (roughly under 30 engineers, single product) | Rarely a deliberate gaming pattern; more often metrics are reported without enough historical baseline to mean much yet | Track trend direction on one or two metrics for internal awareness; avoid presenting DORA benchmark categories to a board as an achievement, since sample size is too small to be meaningful |
| Scale-up (roughly 30–150 engineers, multiple teams, first layer of engineering management) | Cross-team comparison begins informally in leadership syncs; definitional drift starts as individual managers interpret ambiguous rubrics differently | Establish one owner for metric definitions outside line management; stop comparing teams against each other; introduce one counter-metric pairing before adding any new dashboard |
| Enterprise / late-stage (150+ engineers, multiple business units, board and customer-facing reporting) | Definitional gaming becomes structural and can persist for years without detection, because reporting layers separate the board from raw event data by several levels | Full counter-metric pairing across all four core metrics; independent ownership of incident and deployment classification; periodic Metric Integrity Audit as a standing governance process, not a one-time exercise |
What Board Members and Executives Should Ask Before Trusting the Dashboard
Engineering leaders control the eight-step audit above. Board members, CEOs, and other executives who receive delivery metrics secondhand, without direct access to the underlying systems, need a shorter set of pointed questions they can ask in the room, without needing to understand the technical mechanics behind each metric.
"Who decided what counts as a deployment, an incident, and a failure — and has that definition changed recently?" A confident, specific answer naming a cross-functional owner is a good sign. A vague answer, or an admission that the platform team updated the definition last quarter without a formal review, is worth following up on directly.
"What does this number look like without the last definitional change?" If change failure rate improved noticeably in the same quarter a classification rubric changed, ask for the number recalculated under the old rubric. A leadership team confident in its numbers should be able to produce this within a day; one that cannot, or is reluctant to, is signaling something worth a closer look.
"What does customer support say independently?" Ask whether release-linked support ticket volume has been cross-referenced against the internal incident count for the same period, and ask to see both numbers side by side rather than only the engineering-reported one.
"Is any individual's compensation or rating connected to this number, directly or indirectly?" This is the single most diagnostic question in the list. An honest yes is not disqualifying on its own, but it should immediately raise the bar for how much independent verification that specific number receives before it reaches the board.
"How many deployments or incidents is this percentage actually based on?" A change failure rate presented without its underlying sample size can hide the statistical instability described earlier in this article. A rate based on six deployments in a month tells a very different story than the same rate based on two hundred.
"What would it look like if this metric were being gamed right now, and how would we know?" This question is deliberately uncomfortable, and that is exactly why it is useful. A leadership team that has genuinely thought through its own measurement system's failure modes will have a specific, concrete answer. A team that has never considered the question is more likely to be reporting a number nobody has stress-tested.
None of these questions require a board member to understand continuous delivery pipelines or incident-severity rubrics. They require only a willingness to ask who controls a number before deciding how much to trust it.
Frequently Asked Questions
Does this mean DORA metrics are not worth tracking? No. The DORA research remains one of the most rigorously validated bodies of work on software delivery performance available, and the four metrics are genuinely useful proxies for delivery capability when used the way the research intends — as a team's internal trend line against its own history, not as a competitive scorecard or an individual performance input. The argument here is about how the metrics get used, not whether they should exist.
If we already tie deployment frequency to team OKRs, is the damage already done? Not irreversibly, but the incentive to distort the number exists as long as the tie exists. The most direct fix is removing the metric from the OKR and replacing it with a qualitative delivery-capability review that references the metric among other context, rather than treating the number itself as the achievement.
How do we know if a counter-metric itself is being gamed? No single metric, including a counter-metric, is immune to distortion once it carries consequences. The value of pairing comes from using data sources controlled by different teams with different incentives — support tickets are logged by customer success, not engineering, for example — which makes coordinated distortion across both numbers substantially harder, though never theoretically impossible. Treat counter-metrics as a detection tool, not a permanently trustworthy replacement metric. If a counter-metric itself ever starts feeding into a scorecard or a rating, expect it to begin drifting on the same timeline the original metric did, for the same structural reasons — pairing is a detection strategy, not a permanent immunity.
Isn't reclassifying a borderline incident as "not really an incident" sometimes the right call? Sometimes, yes — not every degraded-performance event deserves the same weight as a full outage, and over-classifying minor events dilutes the signal for genuinely serious ones. The problem is not that judgment calls exist; it is that the same party whose performance the classification affects usually should not be the sole authority making that call, and the classification rubric should be reviewed periodically by someone without a stake in keeping the count low.
What is a reasonable first step if we suspect our own dashboard has drifted but don't have time for a full audit? Start with check four from the Metric Integrity Audit above: cross-reference the internal incident log against customer support tickets tagged to recent releases for the last two quarters. It requires no new tooling, usually takes under a day, and is the single check most likely to surface a meaningful gap if one exists.
Does moving to trunk-based development or continuous deployment make gaming less likely? It changes the mechanics but does not remove the incentive. Trunk-based development and small-batch continuous deployment are genuinely associated with better delivery outcomes in DORA's research, but a team under pressure to hit a deployment-frequency target can fragment trivial changes inside a trunk-based workflow just as easily as inside a branch-based one. The practice is not the safeguard; removing the consequence attached to the raw count is.
Should we stop reporting DORA metrics to the board altogether? Not necessarily, but consider reporting them alongside their counter-metrics and with an explicit note on sample size and methodology, rather than as a standalone scorecard. A board that only ever sees a clean quarterly percentage has no way to distinguish a genuinely healthy trend from a number that has quietly been redefined. Boards generally respond well to leadership that volunteers this context unprompted — it reads as rigor, not weakness.
Who inside the organization should own the Metric Integrity Audit once it becomes a standing process rather than a one-time exercise? The most defensible home is a function with visibility into engineering delivery data but no direct stake in any single team's reported numbers — often a QA, reliability, or platform-adjacent role reporting outside the line-management chain being measured. Housing the audit inside the same reporting chain whose performance it evaluates reproduces the exact structural conflict this article describes.
A Restrained Note on Where QAtronic Fits
Detecting metric drift requires cross-referencing data that usually lives in separate systems owned by separate teams — incident logs, deployment pipelines, support tickets, and QA test results rarely speak to each other automatically. QAtronic works with engineering organizations on exactly this kind of cross-system reliability visibility: building the release-validation and incident-classification processes that make change failure rate and time-to-restore numbers trustworthy in the first place, rather than auditing them after the fact. If your delivery dashboard has started to feel disconnected from what your support team or your customers are actually experiencing, that disconnect is usually a process and instrumentation problem before it is a reporting problem, and it is the kind of gap our engineering and QA teams are built to close.
The Distinction Worth Taking Back to Your Team
The question worth asking is not "are our metrics accurate." Every metric described in this article was, in a narrow and defensible sense, accurate at the moment it was reported. The question worth asking is narrower and harder: does anyone in this organization currently benefit, personally or as a team, from one of these numbers moving in a particular direction — and if so, has that person been given quiet authority over how the number gets defined.
A dashboard that stays green while risk quietly compounds underneath it is not a reporting failure in the way a wrong number is a reporting failure. It is a design failure in what the organization chose to reward. The fix is not a stricter audit of last quarter's numbers. It is removing, wherever it currently exists, the connection between a proxy metric and a consequence that gives any individual or team a reason to want that specific number to move regardless of what actually happened underneath it. Take that question back to your own leadership team before the next board deck gets built: which of our four core metrics currently has an owner with something to gain from it looking better than it is, and what would it take to give that authority to someone who doesn't.