Definition of Done Erosion: How "Done" Quietly Becomes "Merged and Demoed"
In the spring of a fintech company's third year, an engineer working on a currency-conversion service asked her lead a reasonable question during a crunch week before a board demo: could the rounding-edge-case test wait until next sprint? The feature worked. She had tested it manually against the three currency pairs the demo would use. Writing the automated test for the fourth decimal-place edge case in a low-volume currency pair would take another half a day, and the pull request already had two approvals waiting on it.
Her lead said yes. Nobody wrote that decision down anywhere except as an unstated understanding between two people in a Slack thread that scrolled out of view within a day. The ticket was closed. It said "Done." Nine months later, a support ticket arrived describing a discrepancy of a few cents in a batch reconciliation report involving exactly that rounding path, in exactly that currency pair. By then, the engineer who had asked the question had moved to a different team, her lead had left the company, and nobody looking at the incident could find any record that the test had ever been scoped, let alone skipped. The team's Definition of Done still listed "new logic covered by automated tests" as a requirement. It had simply stopped being one, for this class of change, for reasons nobody currently at the company could reconstruct.
This is not a story about a careless engineer or a negligent lead. Both made a locally sensible call under real time pressure, the kind every engineering organization makes routinely. The failure was not the decision. The failure was that the decision left no trace, was never revisited, and set a precedent that nobody would recognize as a precedent until a much later incident forced someone to go looking for one.
Why This Isn't a Story About One Bad Call
Most engineering leaders can describe, in vivid detail, a single moment when their team cut a corner under pressure and later regretted it. Fewer can describe the slower, less dramatic process by which a corner became the standard route. That second process is the subject of this article, and it deserves a name of its own: Definition of Done erosion.
A Definition of Done, in its original and still most precise formulation, is what the Scrum Guide calls "a formal description of the state of the Increment when it meets the quality measures required for the product." It exists so that when a team says a piece of work is finished, everyone — including people who did not write the code — can trust that word to mean the same thing every time. The Agile Alliance describes it even more bluntly as a list of criteria the team agrees on and displays prominently, criteria that must be met before an increment counts as done, and notes that work failing to meet it "normally implies that the work should not be counted toward that sprint's velocity." It is, in other words, supposed to have teeth. Failing to meet it is supposed to cost something, visibly, every time.
Erosion is different from a single skipped step because it is cumulative, distributed across many small decisions made by different people at different times, and — this is the part that matters most — never actively chosen by anyone with the authority or the visibility to choose it deliberately. Nobody holds a retrospective and votes to demote "documentation updated" from a requirement to a suggestion. Nobody sends a message announcing that accessibility checks are now optional for internal tooling. It happens the way the fintech rounding test disappeared: quietly, locally, invisibly, one exception at a time, until the aggregate of those exceptions constitutes a new and lower standard that nobody consciously authored and nobody can point to the origin of.
This matters because the failure mode is structurally different from the failure modes engineering leaders are trained to watch for. A team that consciously decides to descope testing for a release under schedule pressure has made a decision that can be logged, revisited, and reversed. A team whose Definition of Done has eroded has not made a decision at all. It has drifted, and drift is much harder to detect, much harder to reverse, and much harder to explain honestly to people outside the team once its consequences surface.
What the Definition of Done Is Actually Supposed to Do
It is worth being precise about what a Definition of Done is for, because the erosion this article describes only makes sense against that backdrop.
A DoD is not a project plan, a sprint goal, or a quality aspiration. It is closer to a contract of ordinary language: an agreement about what the word "done" refers to, so that when a product manager, an executive, a support team, or a customer hears it, they can rely on a fixed and known meaning rather than negotiating that meaning fresh every time. The Scrum Guide's insistence that a Product Backlog item failing to meet the DoD "returns to the Product Backlog for future consideration" rather than being released is the enforcement mechanism that gives the contract its force. Without that consequence, a Definition of Done is a wish list, not a definition.
A well-formed DoD typically bundles several independent quality commitments into one checklist: that new logic is covered by automated tests, that the code has been reviewed by someone other than its author, that user-facing documentation has been updated to reflect the change, that the change has been checked against the product's accessibility commitments (frequently anchored to the Web Content Accessibility Guidelines), that backward compatibility with existing integrations has been confirmed, and that the change has been validated in an environment resembling production closely enough to catch integration failures before customers do. Each of these criteria exists to guard against a specific, identifiable failure: untested logic breaking silently, unreviewed logic containing errors a second set of eyes would have caught, undocumented behavior confusing users and support staff, inaccessible interfaces excluding users who rely on assistive technology, breaking changes disrupting integrations nobody remembered still depended on the old behavior, and staging-only validation missing a production-specific failure.
The instinct to treat these criteria as interchangeable or fungible — as though skipping one is roughly as consequential as skipping another — is itself part of how erosion takes hold. They are not fungible. A skipped documentation update produces a support burden. A skipped accessibility check produces excluded users and, in regulated contexts, legal exposure. A skipped automated test produces a latent defect with an unknown detonation date. Treating "we'll do it next sprint" as a single, general-purpose escape hatch for all of them is precisely the habit that turns individual, defensible exceptions into an undifferentiated pile of unenforced standards.
The Anatomy of an Exception
Erosion has a recognizable shape, and it is worth walking through because the shape is what makes it hard to catch in real time.
It starts, almost without exception, under conditions that are genuinely difficult: a demo commitment made to a board or a major customer, a compounding string of production incidents consuming the team's slack, a departure that leaves a critical review bottlenecked on one remaining senior engineer, or a roadmap commitment made by someone who was not in the room when the DoD was written and does not know it exists. Under those conditions, a specific criterion on the checklist becomes the thing standing between the team and the deadline, and someone with the standing to grant an exception — a lead, a manager, sometimes the individual contributor themselves, operating alone late at night — makes the locally reasonable call to let it slide, just this once.
The phrase "just this once" is worth pausing on, because it is almost always sincere. The person granting the exception is not planning to make a habit of it. They intend to return to the skipped step, and often they genuinely believe they will. What breaks the intention is not dishonesty. It is the absence of any mechanism that would remind them, or anyone else, that the exception exists. The rounding-edge-case test in the fintech example was never going to be forgotten on purpose. It was forgotten because nothing outside two people's memories recorded that it had been deferred rather than completed, and memory is a poor long-term ledger for organizational commitments, particularly once the people involved change teams or leave.
The second exception is easier to grant than the first, for a reason that has nothing to do with the merits of the specific case: the first exception has already established, however quietly, that the criterion is negotiable under pressure. The person granting the second exception may not consciously invoke the first one as precedent, and often the two situations involve different people who never discuss the earlier decision at all. But the underlying fact — that this checklist item bends when the deadline is tight enough — has already entered the team's operating knowledge, even if it never entered its explicit documentation. By the fifth or tenth exception, the criterion has stopped functioning as a requirement in practice, while continuing to exist as a line item in whatever document the team points to when someone from outside asks what "done" means.
The following table lays out how this pattern typically plays out across the criteria that appear in a fairly standard Definition of Done, and what a leader auditing their own team should look for as early evidence of each pattern.
| DoD criterion | Original intent | Common erosion pattern | Early signal to watch for |
|---|---|---|---|
| Automated tests written for new logic | Prevent latent defects from shipping undetected | "We'll add tests next sprint" becomes permanent; tests get written for the easy paths, not the edge cases that prompted the exception | A growing gap between story points closed and test-suite line or branch coverage trend for the same period |
| Code reviewed by someone other than the author | Catch errors, share context, distribute ownership | Reviews become rubber-stamp approvals under time pressure; review turnaround becomes a bottleneck people route around via verbal sign-off | Median time-to-approval dropping sharply while review comment volume also drops |
| Documentation updated for user-facing changes | Keep support, sales, and users aligned with actual behavior | Documentation updates get deferred to "docs week," which rarely arrives; tribal knowledge substitutes for written knowledge | Support tickets asking questions the documentation should already answer |
| Accessibility checked against WCAG-aligned commitments | Ensure the product remains usable with assistive technology | Checks are treated as a launch-blocker only for flagship features, quietly skipped for internal tools, admin panels, and minor UI additions | Accessibility findings clustering in areas built after a specific release, rather than being spread evenly across the product |
| Backward compatibility confirmed | Avoid silently breaking integrations and downstream consumers | "No one uses that old version anymore" becomes an assumption instead of a verified fact | Deprecation notices going unanswered, followed later by an integration partner reporting a break |
| Validated in a production-like environment | Catch integration and configuration failures staging cannot reproduce | Staging drifts from production configuration over time; validation becomes a formality performed against an environment that no longer resembles reality | Incidents whose root cause is described internally as "worked fine in staging" |
None of these erosion patterns require malice, incompetence, or even unusual carelessness. Each is the predictable result of a team under sustained pressure and a checklist with no mechanism for recording, timing, or revisiting the exceptions made against it.
Why Sprint Retrospectives Rarely Catch This
A reasonable objection to everything above is that agile teams already have a built-in mechanism for surfacing exactly this kind of drift: the retrospective. If a team is meant to inspect and adapt its own process regularly, why doesn't Definition of Done erosion show up there before it takes hold?
The honest answer is that retrospectives are structurally biased toward recent, salient, emotionally available events, and DoD erosion is neither recent, salient, nor emotionally available by the time it matters. A retrospective conversation tends to surface what went wrong in the last one or two sprints: a painful incident, a frustrating blocker, a specific pull request that took too long to review. An individual exception granted three sprints ago, resolved without visible consequence at the time, has usually already faded from the discussion by the next retrospective, let alone the fifth one after it. Nobody raises "we skipped the accessibility check on the admin console again" in a retrospective, because in isolation, each individual skip does not feel like an event worth raising. It feels like a minor, forgettable footnote to a sprint that otherwise went fine. The problem only becomes visible in aggregate, and retrospectives, by design, are conversations about the recent past rather than aggregate trend analysis across quarters.
There is a second, quieter reason retrospectives miss this: raising it requires someone to admit, out loud, that they personally granted or accepted an exception, and to do so without any evidence of how common that decision actually is across the rest of the team. An engineer who suspects the documentation criterion has become optional in practice has no way, within the format of a normal retrospective, to know whether they are describing a widespread pattern or an isolated lapse of their own. Raising it risks sounding like either a complaint about a process nobody else seems bothered by, or a confession about a shortcut nobody else is taking. Both readings discourage the observation from being voiced at all, which is exactly how a real, widespread pattern can persist for months without ever surfacing in the forum specifically designed to catch process problems.
This is not an argument against retrospectives, which remain useful for what they are good at: catching acute, recent, sprint-level friction. It is an argument for treating Definition of Done compliance as a measurement problem that needs its own periodic instrument, separate from and complementary to the retrospective, precisely because the retrospective's own structure makes it a poor tool for detecting slow, distributed, aggregate drift. The audit described later in this article exists to fill that specific gap, not to replace the retrospective's role in catching more immediate process friction.
Case One: The Accessibility Criterion That Nobody Killed on Purpose
Consider a hypothetical, but realistic, scenario common to SaaS companies building admin-facing tooling alongside a customer-facing product. This example is illustrative, not a real QAtronic engagement.
A mid-size SaaS company serving mid-market logistics customers had, at its founding, written a Definition of Done that included a specific accessibility criterion: any new interface element affecting keyboard navigation or screen-reader behavior had to be checked before merge. The criterion existed because an early customer, a logistics operator with a legally mandated accessibility requirement for its own workforce tools, had made accessibility a contractual condition of the deal.
Over roughly eighteen months, the product's customer-facing surfaces largely kept honoring the criterion, in part because customer-facing work went through a design review process that happened to include an accessibility pass as a side effect, not because anyone had deliberately preserved the DoD line item there. The internal admin console, used by the company's own support and operations staff rather than by paying customers, told a different story. Three different engineering managers, across three different quarters, made a similar call in similar circumstances: the admin console was "internal only," the checklist item felt disproportionate for a tool used by employees rather than customers, and skipping it bought back a day or two on a release that was already tight. None of the three managers knew that the other two had made the same call. Nothing in the ticketing system distinguished "accessibility check skipped because someone decided it" from "accessibility check performed and passed."
Eighteen months later, the company hired an operations employee who used a screen reader. She found the customer-facing product fully usable and the internal admin console, which she needed for her actual job, substantially broken for her workflow. What surfaced organizationally was not a single failure but a pattern: the admin console had accumulated eleven months of interface changes with effectively zero accessibility verification, none of it logged as an exception, all of it closed in the tracker as "done" against a Definition of Done that still listed the requirement in writing.
The organizational response illustrates the core diagnostic problem this article addresses. When the VP of Engineering asked her team whether the accessibility criterion was being enforced, the honest answer she got back was "usually," which is a non-answer disguised as a partial one. Nobody could tell her, with any precision, what fraction of recent admin-console work had actually been checked, because nobody had ever measured it. The criterion had not been eliminated by decision. It had evaporated by accumulation, and the team's own record-keeping was incapable of proving or disproving the extent of the problem after the fact.
Case Two: The Rounding Test, Revisited as a Systemic Pattern
Return to the fintech example from the opening, but widen the lens from the single incident to what an after-the-fact audit found once the reconciliation discrepancy forced the company to look closely. This scenario, too, is hypothetical and illustrative rather than a documented QAtronic case.
The immediate incident was contained quickly: the rounding logic was corrected, the affected reconciliation reports were reissued, and the financial exposure was small because the currency pair involved low transaction volume. What the postmortem uncovered when it looked past the single defect was less contained. The team pulled forty pull requests closed as "done" over the preceding two quarters that touched the same currency-handling module and checked each one against the automated-test requirement in the team's own written Definition of Done.
Twenty-nine of the forty had a genuine, passing automated test covering the new or changed logic. Six had a test, but one that covered only the primary path and not the edge case the change was actually meant to handle, which the team classified as a test that existed in name but not in substance. Five had no test at all, and none of the five had any comment, ticket field, or commit message indicating that the omission was a deliberate, acknowledged tradeoff rather than an oversight. The team's real compliance rate for the criterion that had failed in the incident was closer to seventy percent than the "essentially always, except in emergencies" belief that had prevailed before the audit. Nobody had lied about this. Nobody had measured it before, either.
The pattern here differs from the accessibility case in an instructive way. In the accessibility example, an entire surface of the product (the admin console) had been carved out of enforcement, a form of erosion that at least followed a somewhat legible boundary. In the fintech example, the erosion was scattered unevenly across a single module, with no discernible boundary at all: some pull requests fully honored the criterion, some partially honored it in a way that looked compliant on a checklist but was not compliant in substance, and some skipped it outright. That inconsistency is itself diagnostic. A criterion still functioning as a real quality gate produces a compliance rate close to the ideal, with departures that are rare, logged, and explainable. A criterion that has partially eroded produces a compliance rate with no clean explanation, scattered across cases that look superficially similar but were handled differently by different people making unrecorded judgment calls.
Case Three: What Happens When You Actually Measure It
A third scenario, again hypothetical, involves a business-to-business software company whose engineering leadership had a specific, nagging worry: that "documentation updated" had quietly become the least-enforced line in their Definition of Done, without anyone being able to say by how much.
Rather than debate the question in the abstract, the VP of Engineering asked a lead on her team to run a structured sample: pull the last sixty tickets marked "done" across three squads, over a period covering roughly ten weeks, and check each one against the documentation-update criterion using whatever evidence existed — a linked documentation pull request, a note in the ticket, a comment from the product manager confirming no user-facing change occurred that would require a doc update. The lead was explicitly asked not to interview anyone about their intentions and not to guess. The rule was: if the evidence doesn't exist, the criterion did not happen, regardless of what anyone remembers believing at the time.
The result: of the sixty tickets, twenty-two involved no user-facing change and were legitimately exempt from the documentation criterion by its own terms; that left thirty-eight tickets where the requirement genuinely applied. Of those thirty-eight, only eleven had a traceable documentation update. Nine had an explicit, logged exception — a comment from the product manager or engineering lead stating that the documentation update was deferred to a specific, named future ticket, with an owner attached. The remaining eighteen had nothing: no update, no logged exception, no comment, nothing indicating the omission had ever been a conscious decision rather than a gap nobody noticed. That is a genuine enforcement rate of roughly twenty-nine percent against the criterion's own applicable population, with an additional twenty-four percent representing legitimate, tracked exceptions rather than failures, and the remaining forty-seven percent representing pure, silent decay.
What made this exercise valuable was not the number itself but what it allowed the VP of Engineering to say to her own leadership with precision: not "we might have a documentation problem," which invites dismissal, but "our Definition of Done requires documentation updates for user-facing changes, and our own recent work shows we are meeting that requirement well under a third of the time it actually applies, with almost half of the gap completely untracked." That distinction between a vague impression and a measured finding is the difference between a conversation that goes nowhere and one that leads to a decision.
The Diagnostic: Auditing Your Own Definition of Done
The three cases above share a structure. In each, the team's actual Definition of Done — the one governing what really shipped — had diverged from its written Definition of Done, and nobody could quantify the gap until someone deliberately measured it. That measurement is the central practical tool this article offers, and it is built to distinguish three categories that get collapsed together far too often in casual conversation about quality:
A healthy criterion is one the sample shows being met at a rate close to what the team would expect from a functioning standard, with departures that are rare, logged, and individually explainable.
A waived criterion, for a specific instance, is one where the requirement was not met, but a record exists — a ticket comment, a linked follow-up item with an owner and a target date, an explicit note from someone with the authority to grant the exception — showing that the gap was a conscious, time-boxed tradeoff rather than an oversight.
A decayed criterion is one where the requirement was not met and no such record exists. The absence of a decision, not the presence of a bad one, is what defines this category, and it is the category that should worry a leadership team the most, because it means the organization has lost the ability to tell the difference between a corner it chose to cut and a corner nobody was watching.
The following framework describes how to run this audit on a real team, in a form specific enough to execute without further design work.
Step one: Pull the team's actual, current Definition of Done, not the aspirational one. Many teams have a DoD written during onboarding or a long-past retrospective that nobody has revisited since. Before auditing compliance, confirm with the team, in a short conversation, which criteria they believe are still active. This step alone sometimes reveals that the team has already informally abandoned a criterion in a way everyone quietly agrees on but nobody has written down; that is a different and less alarming finding than silent, disputed erosion, and the audit should note the distinction.
Step two: Choose a sample window and a sample size that fits the team's own throughput. A reasonable target is forty to eighty tickets or merged pull requests closed as "done" over a period long enough to smooth out single-sprint anomalies, typically eight to twelve weeks. Smaller teams closing fewer than five items a week may need a longer window to reach a meaningful sample; larger teams can use a shorter one. The point is a sample large enough that a handful of well-explained exceptions do not distort the picture, but small enough that one person can review it in a few focused sessions rather than a multi-week project.
Step three: For each criterion in the DoD, define in advance what counts as evidence of compliance. This is the step teams most often skip, and skipping it is what turns an audit into an argument. "Automated tests written" should mean a specific, checkable thing: a test file or test case that did not exist before the change and that exercises the new or changed logic, verifiable by looking at the diff, not by asking the author to recall their intentions. "Documentation updated" should mean a linked documentation change, not a verbal assurance that documentation "probably didn't need updating." Ambiguous criteria produce ambiguous audits; concrete, evidence-based definitions are what make the resulting numbers defensible.
Step four: Review each sampled item against each applicable criterion, recording one of four outcomes: met, waived, decayed, or not applicable. Not applicable matters as its own category, distinct from waived, because inflating the exemption count is one of the easiest ways an audit can quietly flatter a team's real compliance rate. A ticket is not applicable to the documentation criterion only if it genuinely involved no user-facing change; it is not automatically exempt just because nobody remembers whether it needed one.
Step five: Calculate a compliance rate per criterion, not one aggregate score for the whole Definition of Done. A single blended score across all criteria hides exactly the information leadership needs, because it can make a nine-out-of-ten performance on code review and a two-out-of-ten performance on documentation average out to something that sounds acceptable. Report each criterion separately.
Step six: Separate the waived rate from the decayed rate for every criterion, and treat them as fundamentally different findings. A criterion waived eight times out of forty, each time with a named owner and a follow-up ticket, is a team making conscious tradeoffs under real constraints, which is healthy behavior, not a crisis. The same criterion met zero times with zero explanation is an unenforced standard wearing the costume of an enforced one. These two situations look identical on a simple pass/fail chart and require entirely different responses, which is exactly why collapsing them together is the single most common mistake in how teams talk about their own quality.
Step seven: Cross-reference decayed criteria against any incident, defect, or complaint history from the same period. This step is where the audit earns its relevance to leadership rather than remaining an internal engineering curiosity. If a criterion with a high decay rate maps onto an area with a recent, costly failure — as the automated-testing example did in the fintech case — that correlation, stated carefully as a correlation and not overstated as proof, is what turns a compliance number into a business argument.
Step eight: Present findings by criterion, in the format described in the next section, before proposing any remedy. Diagnosing before prescribing sounds obvious, but the temptation to jump straight to "we need a new tool" or "we need stricter process" before establishing what is actually broken produces remedies aimed at the wrong target more often than not.
Step nine: Re-run the same sample structure on a fixed future cadence, not as a one-time event. A single audit tells you where you stand today. A repeated audit, using the same evidence definitions each time, tells you whether a specific intervention actually changed the decay rate or whether the numbers simply moved because a different reviewer applied a slightly different judgment. Quarterly is a reasonable default cadence for most teams; monthly is defensible for a criterion under active remediation.
The following table illustrates what output from this method might look like for a hypothetical mid-size engineering team running the audit against its own five-item Definition of Done for the first time. The figures are illustrative examples built for this article, not real industry benchmarks or the results of any actual QAtronic engagement.
| DoD criterion | Sample size (applicable) | Met (healthy) | Waived (logged exception) | Decayed (silent) |
|---|---|---|---|---|
| Automated tests for new logic | 52 | 71% | 12% | 17% |
| Code review by a second engineer | 58 | 91% | 3% | 6% |
| Documentation updated | 38 | 29% | 24% | 47% |
| Accessibility check completed | 21 | 33% | 5% | 62% |
| Backward compatibility confirmed | 30 | 63% | 10% | 27% |
What a table like this makes visible, at a glance, is that a team's Definition of Done rarely erodes uniformly. Code review, in this illustrative example, remains genuinely healthy, likely because the pull-request workflow makes skipping it structurally awkward rather than because anyone consciously prioritized it. Documentation and accessibility, by contrast, show decay rates high enough that calling them "requirements" is closer to fiction than description. This unevenness is itself useful information: it tells leadership exactly where to direct attention, rather than treating "quality" as one undifferentiated problem requiring one undifferentiated fix.
Presenting the Findings Without It Reading as Blame
An audit like this can go one of two ways once it reaches leadership. It can land as a diagnostic that improves the organization's ability to see itself clearly, or it can land as an accusation that makes individual engineers and managers defensive, incentivizes them to make the next audit less honest, and ultimately makes the underlying erosion worse by driving it further underground. The difference is almost entirely a matter of framing, and it deserves as much deliberate attention as the audit method itself.
The framing that works starts from a premise that happens to be true in the overwhelming majority of real cases: nobody decided to lower the standard. The audit's entire value proposition rests on this fact, and stating it explicitly, early, and sincerely defuses most of the defensiveness before it starts. "We looked at our last quarter of closed work against our own Definition of Done, and here's what we found" is a very different opening than "we found people skipping tests," even when the underlying data is identical. The first framing locates the finding at the level of the system — the checklist, the workflow, the incentives, the absence of a tracking mechanism — and the second locates it at the level of individual behavior, which is both less accurate, given how the erosion actually happened, and more likely to produce a defensive, unproductive reaction.
It also helps to present the waived and decayed categories together as evidence that the team is capable of honest, disciplined tradeoffs, not just as evidence of a problem. A criterion with a healthy waived rate and a low decayed rate is not a failure to hide; it is a team making visible, accountable calls under real constraints, which is exactly the behavior a good Definition of Done process should produce. Leading with that example, where one exists in the data, establishes that the goal of the exercise is not zero exceptions — an unrealistic and actually undesirable target — but zero invisible exceptions.
When presenting a high decay rate on a specific criterion, resist the urge to identify which individuals were responsible for which instances unless the audit is being used for a narrowly scoped, already-agreed process change rather than a general status update. The goal of a status conversation is to establish a shared, accurate picture of where the standard currently stands and to agree on what should happen next, not to assign personal culpability for a pattern that, as this article has argued from the outset, virtually never traces back to a single deliberate choice by a single person. Naming individuals in that setting shifts the conversation from "how do we fix the system" to "how do I defend myself," and the second conversation produces worse decisions and worse data next quarter, because people will quietly route around whatever mechanism generated the audit.
Finally, come to the conversation with a specific, proportionate proposal rather than an open-ended alarm. "Our documentation criterion has decayed to a twenty-nine percent enforcement rate, and we're proposing to add a required documentation-link field to our ticket template, re-measure in one quarter, and treat repeated undocumented exceptions the same way we'd treat a repeated missed test" is concrete, bounded, and testable. "Our documentation is a mess and we need everyone to care more" is not a proposal at all; it is a mood, and moods do not survive contact with the next deadline any better than the original, unenforced checklist item did.
Turning Exceptions Into a Countable Category of Technical Debt
The audit method above solves the detection problem. Solving the recurrence problem requires a change to how the team talks about and tracks exceptions going forward, not just a one-time cleanup.
Martin Fowler's technical debt metaphor is useful here in a way that goes beyond its common, looser usage as a synonym for "code we don't like." Fowler's original formulation distinguishes deliberate debt, taken on knowingly for a reason, from inadvertent debt, which accumulates because a team was not paying attention to consequences it did not anticipate. He is explicit that inadvertent debt is frequently the more dangerous kind, precisely because nobody chose it and therefore nobody is positioned to manage its repayment. A Definition of Done exception is a perfect, almost textbook instance of this distinction, and treating it through the same lens Fowler applies to code quality more broadly clarifies what the fix actually needs to accomplish.
A genuine, deliberate DoD exception should be logged with the same rigor a team would apply to any other explicitly acknowledged technical debt item: what was skipped, why, who approved the exception, and — critically — a specific, named point at which it will be revisited, whether that is a linked follow-up ticket, a target sprint, or a triggering condition ("revisit before this endpoint handles production traffic above one thousand requests per minute"). This is not bureaucratic overhead for its own sake. It is the mechanism that converts an exception from something that quietly becomes the new normal into something that stays visible until it is either resolved or consciously re-approved.
A criterion that shows up in the audit as decayed rather than waived should be treated as a distinct finding, one that indicates the checklist item itself, not any single piece of work, needs organizational attention. The appropriate response to a high decay rate is rarely "everyone must now follow the rule perfectly." It is closer to: why has this criterion become the one people quietly bypass under pressure, and is the answer that the criterion is genuinely important and needs a structural change to make honoring it easier (for instance, a documentation-link field required before a ticket can be closed, or a lightweight automated accessibility check integrated into the pull-request pipeline rather than a manual step someone has to remember), or is the answer that the criterion has outlived its usefulness and should be explicitly renegotiated rather than left to erode by accident.
This distinction matters because not every DoD criterion deserves to survive forever in its original form, and pretending otherwise creates its own kind of dishonesty.
Building the Exception-Logging Habit Into Tools You Already Use
None of this requires new software. It requires a small, specific change to fields and templates most teams already maintain, applied consistently enough that logging an exception becomes the path of least resistance rather than an extra chore competing with the deadline that prompted the exception in the first place.
The simplest version is a required field on the ticket or pull-request template itself, positioned next to wherever the team already marks a criterion as complete. Rather than a single checkbox labeled "tests written," a more honest template offers three states for each criterion: met, waived (with two required sub-fields — reason and revisit date or owner), and not applicable (with a one-line reason, since "not applicable" is itself a claim that deserves a trace, not a silent default). A pull-request description template built around this structure might look like the following, adapted to whatever criteria a specific team's Definition of Done actually contains:
## Definition of Done
- Automated tests for new logic: [met / waived / not applicable]
- If waived or not applicable, reason:
- If waived, revisit by (ticket link or date):
- Code reviewed by a second engineer: [met / waived / not applicable]
- Documentation updated: [met / waived / not applicable]
- If waived or not applicable, reason:
- Accessibility check completed: [met / waived / not applicable]
- Backward compatibility confirmed: [met / waived / not applicable]
The value of this structure is not that it prevents anyone from skipping a criterion. It does not, and it is not meant to. Its value is that it makes skipping a criterion require one deliberate, typed sentence instead of zero, which is precisely the gap that separates a waived exception from a decayed one. A reviewer approving the pull request sees the claim and can push back on it in the moment, which is the cheapest point in the entire process to catch a questionable exception, far cheaper than catching it nine months later in an incident review.
Where a team has the engineering capacity to go further, a lightweight automated check — a continuous-integration script that fails the build, or at minimum posts a visible warning, if a pull request touching source files contains no corresponding change to a test file, and the "waived" field is empty — closes the remaining gap between a template that asks for honesty and a mechanism that verifies it. This does not need to be sophisticated. A rule as blunt as "if this diff modifies files under src/ and does not modify any file under test/ or matching *.spec.*, require the waived field to be filled in or fail the check" catches a meaningful share of unlogged test-writing exceptions without requiring any deep static analysis, and most CI systems can express a rule at that level of bluntness in a few dozen lines of configuration.
The underlying principle generalizes past testing specifically: wherever a Definition of Done criterion has a corresponding artifact that should exist when the criterion is met — a test file, a documentation page diff, an accessibility-scan report attached to the ticket — a lightweight, low-precision automated check for the artifact's absence is nearly always cheaper to build and maintain than the manual audit described earlier in this article, and the two are complementary rather than redundant. The automated check catches missing artifacts continuously and immediately; the periodic sample audit catches the subtler case where an artifact exists but is hollow, such as the fintech example's test file that covered only the primary path rather than the edge case the change was actually meant to address, which no simple presence check would flag but a human reviewing the sample would.
When a Criterion Should Be Retired on Purpose
There is a genuine difference between erosion and evolution, and a mature engineering organization needs language for both, because treating every decayed criterion as a crisis to be restored to full enforcement is its own mistake.
Some Definition of Done criteria are written in response to a specific, time-bound circumstance — a single demanding customer's contractual requirement, a compliance regime the product no longer falls under after a business-model change, a technology dependency the team has since replaced — and outlive the reason they were added. When that happens, the healthy response is not silent abandonment, which is exactly the failure mode this article has been describing, but an explicit decision, made by whoever owns the Definition of Done, to remove or rewrite the criterion, documented with the same clarity as its original adoption. The team should be able to point to when and why "check for compatibility with the legacy on-premises deployment" left the checklist, in the same way they should be able to point to when and why it was added.
The practical test for distinguishing a criterion that deserves retirement from one that is simply being allowed to decay is whether the removal is something the team would be comfortable stating out loud, in a retrospective or a written update, without embarrassment. "We removed the on-premises compatibility check because we discontinued on-premises deployments eight months ago and the criterion had become meaningless" is a defensible, healthy statement. "We stopped doing accessibility checks on the admin console because it felt disproportionate for internal tooling" is a statement that, if a team is honest with itself, probably should have been raised and debated explicitly before it became a pattern, precisely because it embeds a judgment about whose usability matters that deserves more scrutiny than a quiet, distributed non-decision gives it.
A useful discipline, once a team starts running the audit described above, is to require that any criterion showing a decay rate above a threshold the team sets for itself — something in the range of a third to a half of applicable cases going unmet without a logged exception is a reasonable starting point, adjusted to the criterion's actual stakes — triggers one of exactly two outcomes within a defined period: either a structural change that makes the criterion genuinely easier to honor, or an explicit, documented decision to retire or rewrite it. What a decayed criterion should never be allowed to do is simply continue decaying indefinitely while remaining on the books as though it were still a live requirement, because that is precisely the condition that produces the gap between a written standard and a real one that this entire article is about.
Differences Across Startups, Scale-Ups, and Enterprises
The mechanics of erosion are broadly similar across company sizes, but the conditions that produce it, and the remedies that fit, differ enough to be worth separating.
Early-stage startups often have a Definition of Done that is genuinely minimal by design, sometimes deliberately so, because the cost of over-specifying quality gates before product-market fit is real and the argument for moving fast is not merely a rationalization. The risk at this stage is less that a rich DoD erodes and more that the team never establishes enough of one to erode, which produces a different but related failure: nobody can distinguish a deliberate, appropriate minimalism from an accidental one, because the question was never explicitly decided in the first place. The audit method above still applies, just against a shorter list of criteria, and the value at this stage is establishing, with evidence, whether the team's actual practice matches what its founders believe it to be.
Scale-up companies, roughly the size range where a Definition of Done was written by an early team that has since partially turned over, are where the fintech and SaaS examples in this article are most typical. The original context for each criterion has faded from institutional memory, new engineers inherit a checklist without necessarily understanding which items were hard-won responses to a real past failure and which were aspirational additions nobody has ever actually tested, and the volume of work has grown past the point where a small, tight-knit team's informal, memory-based tracking of exceptions can keep up. This is the stage at which formalizing the exception-logging habit described above tends to produce the largest return, because it is also the stage at which the informal version of that habit is most likely to be silently failing.
Enterprises with mature, heavily documented Definition of Done processes face a different, subtler version of the same problem: the checklist itself can become so long and so detached from any single team's daily reality that compliance becomes theatrical rather than substantive, with teams learning to check boxes in a tracking system without the underlying evidence the box is supposed to represent ever existing. In this setting, the audit's insistence on evidence-based, not self-reported, compliance checking is especially important, because self-reported compliance in a large organization tends to converge toward whatever the tracking system rewards, regardless of what actually happened in the code.
Where This Connects to Delivery Stability, Not Just Individual Incidents
It would be a mistake to treat Definition of Done erosion purely as a story about isolated incidents like the fintech rounding bug or the SaaS accessibility gap, satisfying as those individual stories are for illustrating the mechanism. The pattern also shows up at the level of aggregate delivery performance, which is where research on software delivery more broadly becomes relevant.
The DORA research program's 2024 State of DevOps report, produced by Google Cloud's DevOps Research and Assessment team, found that increased adoption of AI-assisted development tools correlated with measurable improvements in some quality-adjacent measures, including code quality and documentation quality as self-reported by respondents, while simultaneously correlating with a decrease in software delivery stability and throughput over the same period. The report's own framing of this tension is instructive: it states plainly that "improving the development process does not automatically improve software delivery — at least not without proper adherence to the basics of successful software delivery, like small batch sizes and robust testing mechanisms." That finding, while specifically about AI tool adoption, generalizes to the broader argument of this article. A tool, a process change, or a team's stated intentions do not, by themselves, guarantee that the underlying quality contract is being honored. Only the underlying discipline of consistently applying fundamentals like testing does that, and the DORA report's own data suggests organizations frequently assume the former substitutes for the latter when it does not.
This is worth connecting explicitly to Definition of Done erosion because the two dynamics compound each other in a specific way. A team whose actual test-writing discipline has quietly decayed, in the manner described throughout this article, is a team whose delivery stability is already more fragile than its dashboards suggest, because the dashboards typically measure throughput and velocity, not the compliance rate of the quality gates that are supposed to make that throughput sustainable. Layering a new tool, a faster release cadence, or an ambitious roadmap commitment onto that already-eroded foundation does not create a new problem so much as it accelerates the arrival of the incident that the erosion was always going to eventually produce.
A Simple Cost Comparison: Logging an Exception Versus Discovering It Later
Leadership conversations about process changes tend to stall on an unspoken cost question that is worth making explicit: does the discipline this article recommends cost more than the problem it prevents? The honest answer requires comparing two different kinds of cost, one small and immediate, the other larger and deferred, using illustrative figures rather than claimed industry benchmarks, since no verified public dataset measures this specific comparison across organizations.
The immediate cost of logging a genuine exception, using the template described above, is close to the time it takes to type one or two sentences into a field that already exists in the pull-request description: a minute or two per instance, occurring at the moment the exception is granted, when the context is freshest and cheapest to record. The immediate cost of the periodic sample audit, run quarterly against a sample of forty to eighty tickets as described earlier, is a few focused hours for one experienced reviewer, several times a year.
The deferred cost of an undetected, decayed criterion is harder to bound precisely, because it depends on which criterion decayed and what eventually surfaces the gap, but the fintech and SaaS examples earlier in this article illustrate its shape: an incident response and remediation effort, a support and reissue process for affected customers, a postmortem consuming multiple engineers' and managers' time to reconstruct a decision nobody logged in the first place, and, in the accessibility case, a potential compliance and reputational exposure that a logged, time-boxed exception would never have created, because a logged exception would have forced someone to weigh that exposure explicitly at the time rather than never at all.
The table below lays out this comparison in illustrative, order-of-magnitude terms, not as a validated benchmark but as a way of making an otherwise abstract tradeoff concrete enough to discuss with a leadership team that is, reasonably, asking whether any of this is worth the friction.
| Activity | Approximate time cost | When it is paid | Who pays it |
|---|---|---|---|
| Logging one exception in the PR template | One to two minutes | At the moment the exception is granted | The engineer or lead making the call |
| Reviewer pushing back on a questionable logged exception | A few minutes of discussion | During code review, before merge | Reviewer and author together |
| Running one quarterly sample audit (40–80 tickets) | A few hours | Once per quarter, on a predictable schedule | One reviewer, typically a lead or QA-focused engineer |
| Responding to an incident traced to an undetected, decayed criterion | Days to weeks, spread across response, remediation, and postmortem | Unpredictably, often months after the originating exception | Multiple engineers, a manager, sometimes a support or legal function |
| Reconstructing what was actually decided, after the people involved have moved on | Often impossible to complete accurately at all | During the postmortem for the incident above | Whoever is assigned to the postmortem, working with incomplete information |
The asymmetry in that table is the entire economic argument for treating exceptions as a countable category rather than an unspoken habit. The cost of visibility is measured in minutes and scheduled hours. The cost of invisibility is measured in unscheduled days, and it is frequently paid, in part, by people who were not present for the original decision and have no way to make it fully legible even when they are specifically tasked with doing so.
Questions to Take Into Your Next Leadership Review
A leader convinced by the argument above still needs a way to raise it upward without simply forwarding this article. The following questions are specific enough to prompt a real answer rather than a reassuring generality, and they are ordered from least to most uncomfortable, which is usually the right order to ask them in.
What does our Definition of Done currently say, in writing, and when was it last reviewed by the people actually held to it. Many leadership teams discover, on asking this question directly, that the written document and the team's own understanding of it have already diverged, independent of any compliance question.
If we sampled forty of our own closed tickets from the last quarter against that written document, what fraction would show real, traceable evidence for each criterion, as opposed to a belief that the criterion was probably met. This is the question that the audit method in this article is built to answer, and asking it before commissioning the audit is a useful test of whether the organization already suspects the answer will be uncomfortable.
Of the exceptions we already know about, how many have a named owner and a specific date or trigger for revisiting them, versus how many exist only as something someone remembers deciding. This question tends to surface the gap between the exceptions a team is proud of, because they were handled deliberately, and the ones nobody would volunteer, because they were not.
Which of our current criteria would we be comfortable retiring outright, in writing, if we admitted it no longer reflects a standard we actually intend to hold ourselves to, rather than letting it continue existing on paper while quietly not existing in practice. This question separates the honest case for evolving a standard from the dishonest case for simply not enforcing it.
What would it have cost us, in the last incident or near-incident we can point to, if the decision that led to it had been logged, visible, and revisited on schedule rather than made once and forgotten. This last question is the one that usually moves a leadership conversation from abstract process discussion to a concrete decision to fund the audit, because it ties the entire argument back to a cost the organization has already paid at least once.
A Short Diagnostic Checklist for a First Read of Your Own Team
For a leader who wants a faster, lower-effort first signal before committing to the full nine-step audit above, the following questions can be answered in a single working session and will usually indicate whether a fuller audit is warranted:
- Can anyone on the team name the last three explicit, logged exceptions granted against the Definition of Done, including who approved each one and when it was supposed to be revisited?
- If you pulled ten recently closed tickets right now, could you find documented evidence — not a memory, an actual artifact — that each Definition of Done criterion was met, for each ticket, in under two minutes per ticket?
- Does your team's ticket template or pull-request template have a field that captures a DoD exception, or does an exception currently live only in a Slack message, a verbal conversation, or someone's memory?
- Has any criterion on your Definition of Done gone unmentioned in a retrospective, a review, or a leadership conversation for longer than two full quarters?
- If a new engineer joined the team tomorrow and read your written Definition of Done literally, would experienced team members immediately say "well, we don't really enforce that one" about any item on the list?
A "no" to the second question, or an immediate, unhesitating "yes" to the fifth, is usually sufficient reason to run the full audit rather than treat the quick check as reassuring on its own. This shorter version is a smoke detector, not a replacement for the inspection; it tells you whether to look further, not what you will find when you do.
Frequently Asked Questions
Is a Definition of Done the same thing as "acceptance criteria"? No, and conflating the two is a common source of confusion. Acceptance criteria are specific to a single piece of work and describe what that particular feature or story must do to satisfy its intended behavior. A Definition of Done is a standing, team-wide standard applied to every piece of work regardless of its specific content — the same automated-testing, review, and documentation requirements apply whether the ticket is a new feature or a one-line bug fix. A story can fully satisfy its acceptance criteria while still failing its Definition of Done, if, for instance, the feature behaves exactly as specified but was never covered by an automated test.
What is the difference between "done" and "done done"? The informal phrase "done done" emerged in agile teams specifically to name the gap this article is about: work that is functionally complete and demoable, versus work that has actually satisfied every criterion in the team's Definition of Done, including the less visible ones like documentation and test coverage. Teams that need the phrase "done done" to distinguish these two states have usually already, if unintentionally, admitted that their Definition of Done has partially eroded into an aspirational document rather than an enforced one; a team with a genuinely enforced DoD should not need a second, informal term to describe what "done" was already supposed to mean.
How is this different from just having weak QA or insufficient test coverage? Weak QA or low test coverage can result from many causes, including an honest, resourcing-driven decision never to build a strong automated-testing discipline in the first place. Definition of Done erosion specifically describes a situation where a standard did exist, was written down, and was at some point genuinely enforced, and has since been quietly abandoned in practice without anyone deciding to abandon it. The distinction matters because the fix is different: a team that never had adequate test coverage needs investment and a plan; a team whose test-coverage requirement has eroded needs visibility into what changed and why, since it may already have the skills and tooling to meet the standard and simply stopped being held to it.
Doesn't logging every exception just create bureaucracy that slows the team down further? Logging a genuine exception takes one sentence in an existing ticket field or comment: what was skipped, who approved it, and when it will be revisited. That is a materially smaller burden than the conversation, incident response, and postmortem that typically follow an undocumented exception once its consequences surface, sometimes months later, once the people involved have moved on and the context has to be reconstructed from scratch. The goal is not to add process for its own sake; it is to make an already-happening decision visible at the moment it happens, rather than invisible until it fails.
Should every single Definition of Done criterion be audited this way, or just the ones we're worried about? Auditing the full set at least once is worth the modest additional effort, because the criteria a team is not worried about are sometimes the ones with the highest hidden decay rate, precisely because nobody has been watching them. The code-review criterion in the illustrative table above showed strong health partly because the pull-request workflow made skipping it structurally awkward; a criterion without an equivalent structural nudge, like documentation, is exactly the kind that decays without anyone noticing until it is measured.
What should we do if the audit shows a criterion has decayed almost completely? Treat the finding as information about the system, not as evidence to relitigate individual decisions made months or years ago by people who may no longer be on the team. The two constructive paths forward are making the criterion genuinely easier to honor through a structural change (an automated check, a required field, a lighter-weight version of the requirement that still serves its original purpose) or explicitly, deliberately retiring or rewriting the criterion if it no longer reflects a standard the team actually intends to hold itself to. What should not happen is leaving the criterion on the books, unenforced and unaddressed, for another audit cycle.
Who should own running this audit — engineering, QA, or an outside reviewer? Any of the three can run it competently, provided the person doing it is willing to apply the evidence rule strictly, including to work their own team produced. In practice, an internal reviewer sometimes struggles with exactly that impartiality, particularly when the sample includes their own recent pull requests or their direct reports' work, which is part of why some organizations prefer a QA lead outside the immediate team, or an external reviewer, for at least the first pass. What matters far more than who runs it is that the evidence definitions from step three are fixed in advance and applied identically regardless of whose work is being sampled, since an audit that quietly grades familiar names more generously than unfamiliar ones will produce numbers nobody should trust, including the numbers that happen to look reassuring.
The Distinction That Matters Going Forward
The argument of this article rests on a single distinction, and it is worth stating it once more, plainly, as the thing to carry back to a team rather than as a summary of everything covered above: a corner that gets cut with a name, a reason, and a date attached is a tradeoff. A corner that gets cut with nothing attached to it is a debt nobody knows they are carrying, and debts nobody knows they are carrying are the ones that come due at the worst possible time, traced back by an incident review to a decision nobody remembers making.
The question worth taking back to an engineering team is not "are we meeting our Definition of Done." Almost every team will answer yes to that question, sincerely, because the belief that the DoD is being honored survives long after the practice of honoring it has quietly stopped. The more useful question, and the one this article's diagnostic is built to answer with evidence rather than impression, is: for each criterion on our own checklist, if we pulled the evidence right now, would it agree with us?
Where QAtronic Fits
Running the audit described in this article, honestly and on a recurring cadence, requires exactly the kind of structured, evidence-based review that an internal team, understandably consumed by its own delivery pressure, often struggles to prioritize for itself. QAtronic works with engineering teams as a software testing services partner on precisely this kind of independent quality assessment: reviewing a sample of recently closed work against a team's own stated standards, separating genuinely enforced criteria from quietly abandoned ones, and helping translate the findings into a specific, prioritized remediation plan rather than a vague call to try harder. If your team suspects its Definition of Done has drifted from what it once meant but has never measured the gap, that is a scoped, well-bounded engagement, not an open-ended audit of everything your organization does.
Resources and Sources
- The 2020 Scrum Guide — Scrum.org / Ken Schwaber and Jeff Sutherland
- Definition of Done — Agile Alliance
- bliki: TechnicalDebt — Martin Fowler
- Accelerate State of DevOps Report 2024 — DORA / Google Cloud
- Web Content Accessibility Guidelines (WCAG) 2.2 — World Wide Web Consortium (W3C)
- Mid-Sprint Reprioritization: The Hidden Quality Cost — QAtronic
- Scope Creep Risk Management: A Framework for TPMs — QAtronic