Software Estimation Accuracy: Why Estimates Always Miss
Share this post

Why Software Estimation Accuracy Breaks in the Same Direction Every Time

The postmortem for a slipped fintech release usually reads the same way. The team lists what went wrong: a payment processor's sandbox behaved differently than production, a compliance review surfaced a new requirement in week five, a senior engineer went out on leave during the critical path. Everyone nods. These are real events, and they really did contribute to the delay. The retro closes with action items about better documentation and earlier stakeholder involvement, and the team moves on to the next sprint carrying the same estimation process into the same kind of commitment.

What almost never appears in that postmortem is the estimation meeting that happened eight weeks earlier, where an engineer said "probably six weeks" in response to a product manager's question, a VP wrote "6 weeks" into a roadmap slide before the sentence was finished, and everyone in the room subsequently treated that number as a finding rather than a guess. Nobody in that room did anything indefensible. The number felt reasonable. It was based on genuine technical judgment about how long the coding would take. It was wrong in a way that had almost nothing to do with the sandbox environment, the compliance surprise, or the medical leave, and almost everything to do with what "six weeks" was quietly assumed to include.

This is the argument this article makes: the interesting failure in most late, bug-heavy releases isn't the list of things that went wrong during the sprint. It's the estimate itself, and it's wrong in a predictable direction, for reasons that are now well studied outside of software engineering and increasingly well documented inside it. Testing, integration, and edge-case handling are the categories of work most reliably left out of the number people agree to at the start. When the schedule later gets defended and the scope doesn't shrink, those are exactly the categories that get compressed first, silently, without anyone deciding to cut them. "We shipped late" and "we shipped with more defects than usual" are frequently not two separate incidents. They are the same estimation failure, observed from two different points on the calendar.

This matters for how engineering and product leaders should run planning, and it matters more than most retrospectives acknowledge. A team that keeps fixing the symptom (more process rigor during the sprint, more QA headcount, more test automation investment) while leaving the estimation process untouched will keep producing the same outcome under a different name. The fix has to start earlier, at the point where the number gets said out loud.

The Postmortem Looks at the Wrong Meeting

Most engineering organizations have gotten reasonably disciplined about examining what happens after a commitment is made. Sprint retrospectives, incident reviews, and release postmortems are standard practice at any company past its earliest stage. They ask good questions: What blocked us? What did we not anticipate? What should the runbook have caught?

Almost none of that machinery looks backward past the start of the sprint to the estimation session itself, because the estimation session doesn't produce an artifact that invites scrutiny the way an incident does. Nobody pages anyone when an estimate is wrong. There's no on-call rotation for a planning meeting. The consequence of the estimate arrives weeks later, disguised as a scheduling problem, a scope problem, or a testing problem, by which point it looks like an unrelated event with its own unrelated causes.

Consider what actually happens between a plausible-sounding estimate and a rough release, in the ordinary case where nothing dramatic goes wrong.

An engineering team estimates a feature at three weeks. The estimate is built, implicitly, around "how long will it take me to write the code." Nobody phrased it that way in the meeting; nobody would defend that framing if asked directly. But when three experienced engineers separately try to picture the work, they picture themselves opening an editor and building the thing, because that's the part of the job that's vivid to imagine. Testing the feature against real data, wiring it into three existing services, handling the account states that don't fit the common case, and fixing whatever the first round of testing turns up are technically part of "the work," but they aren't in the mental image anyone builds when they say a number.

The team spends three weeks coding. This part goes close to plan, because coding time is the part that got estimated. Now integration testing starts, and it turns up: an assumption about currency formatting that breaks for one of the supported markets, a race condition that only appears under concurrent requests, and a third-party API that silently changed its rate-limit behavior since the last time anyone touched that integration. None of this is unusual. It's what integration testing is for. But there was no time allocated for it, because the three weeks are gone and the release date was set from those three weeks. The team now has two options that both look like separate problems from the outside: slip the date (a "delivery" failure) or ship without fully resolving what testing surfaced (a "quality" failure). Frequently, under real deadline pressure, both happen at once: a shorter slip plus a riskier release than the team would have chosen with more room.

The retrospective that follows this release will accurately describe the currency bug, the race condition, and the API change as proximate causes. It will be technically correct and almost entirely beside the point. The proximate causes are what testing time is supposed to find and what estimation time is supposed to budget for. The actual failure happened when nobody put a number next to "time to find and fix the things testing will find," because that time is invisible until testing happens, and estimates are built almost entirely from what's visible in advance.

This is not a story about a careless team. It's the default behavior of estimation under specific, identifiable cognitive conditions, and understanding those conditions is the difference between fixing the process and fixing the symptom.

The Research Behind Why Estimates Miss in a Predictable Direction

The tendency to underestimate how long a task will take, even when a person has direct experience with similar tasks, is not folklore. It has a name and a research history that predates modern software engineering by decades.

The planning fallacy: an inside-view problem

Daniel Kahneman and Amos Tversky described what they later called the planning fallacy in their work on intuitive prediction in the late 1970s, and Kahneman revisited and extended the concept with Dan Lovallo in later work on decision-making under uncertainty (Planning fallacy — Wikipedia). The core finding, replicated across many contexts since, is that when people forecast how long a task will take or how much it will cost, they tend to focus on the specific plan in front of them (the steps they can currently picture) rather than on how similar plans have actually turned out in the past. Kahneman and Tversky called this the "inside view."

The inside view feels rigorous because it involves real analysis. A team walks through the actual feature, the actual code paths, the actual known dependencies. The problem is that the inside view can only account for what the estimator currently knows to picture. It cannot account for the unknown unknowns (the integration surprises, the requirement that turns out to be ambiguous, the edge case nobody thought to ask about) because by definition those aren't part of anyone's mental model yet. The inside view systematically excludes the categories of work that are hardest to picture in advance, and testing is disproportionately the discipline whose entire purpose is finding exactly that category of problem. This is not a minor detail. It is close to the mechanical explanation for why testing and integration time are the most reliably underestimated categories of software work: they are, by design, the phase where the unknown unknowns get converted into known ones, and the inside view has no place to put them until they've already appeared.

The alternative Kahneman and Tversky proposed is the "outside view": instead of reasoning from the details of this specific plan, look at a reference class of similar past efforts and use their actual outcomes as the starting point for a forecast, adjusting only for genuine, evidenced differences. The outside view is less intuitively satisfying (it doesn't feel like it uses what the team specifically knows about this feature) but it is the more accurate predictor precisely because it isn't limited by what the team can currently imagine going wrong.

  Inside view (how most estimates get built) Outside view (reference-class forecasting)
Starting point The specific plan, broken into the steps the team can currently picture A set of comparable past projects or tickets and how they actually turned out
What it captures well Known technical steps, familiar code paths, tasks the team has explicitly discussed Unknown unknowns, categories of work that never make it into the plan discussion, systemic patterns across many efforts
Where it fails Cannot budget for problems nobody has thought to picture yet; almost by definition, this includes most integration and edge-case work Requires a real base rate; weak or too-narrow a reference class produces a weak forecast
Typical bias direction Optimistic; treats the visible, plannable work as the whole of the work Tends to be more calibrated because it is anchored to what actually happened before, not what seems plannable now
Effort to produce Feels rigorous, is fast to generate, matches how planning meetings are structured Requires tracked historical data and a deliberate step to consult it
Common failure mode when unmanaged Coding-time-as-proxy-for-done: the visible, codeable steps stand in for the whole deliverable Reference class chosen too narrowly, or adjusted back toward the inside-view number under social pressure

Bent Flyvbjerg, a professor best known for his empirical work on megaproject cost and schedule overruns, extended reference-class forecasting from a psychological finding into a practical forecasting method, and it has since been adopted by government transportation and infrastructure agencies as a mandated part of cost estimation for major public projects (Reference Class Forecasting in 90 Seconds — Major Projects Association; Underestimating Costs in Public Works Projects: Error or Lie? — World Bank). Flyvbjerg's central empirical contribution is that cost and schedule overruns on large projects are not randomly distributed around an accurate mean. They are skewed almost entirely toward underestimation, across an enormous number of projects, across decades, across countries and sectors that have little in common except that humans built the forecasts using the inside view. A distribution that is skewed in one consistent direction is not describing noise. It's describing a bias, and a bias is a thing you can structurally correct for, which is a fundamentally different problem than random error, which you can only average out.

This distinction is the foundation of the argument in this article. If software estimates were simply noisy (sometimes too high, sometimes too low, averaging out to roughly accurate) the fix would be statistical: add a margin, round up, track a standard deviation. Software estimates are not noisy in that sense. They are biased in a specific direction, and specifically biased against the categories of work that are least visible at the moment the estimate gets made. Padding a systematically biased number, as the next section covers, does not fix the bias. It just moves the argument to a different number.

What the software-specific research adds

Kahneman, Tversky, and Flyvbjerg were not writing about software, and it would be a mistake to import their findings as though software estimation were identical to public infrastructure forecasting. Fortunately, software engineering has its own research tradition on effort estimation, running for more than two decades, that reaches compatible conclusions using project data instead of psychology-lab or infrastructure data.

Magne Jørgensen, a Norwegian researcher who has spent much of his career studying how software professionals actually estimate, published a widely cited review of decades of expert-estimation studies in the Journal of Systems and Software (A Review of Studies on Expert Estimation of Software Development Effort — Jørgensen). The review's broad finding is consistent with the planning fallacy framing: expert, intuition-based estimates in software development are frequently optimistic, are strongly anchored by information presented before the estimate (including numbers suggested by whoever is asking for the estimate), and are not reliably improved simply by being made by more senior or more experienced engineers. Seniority helps with technical judgment about how to build something. It does not automatically correct for a structural bias in how the estimating process treats invisible work, because the bias isn't a knowledge gap: it's a property of how the inside view constructs a plan.

This last point deserves emphasis because it's frequently misunderstood inside engineering organizations. When a schedule slips, a common internal narrative is that a specific engineer or team "isn't good at estimating," implying that better individuals, or more experience, would fix the problem. The research doesn't support that read. Estimation accuracy is largely a property of the process: what information estimators are given, what categories of work they're explicitly asked to itemize, whether they consult historical outcomes or only the current plan, and what social pressure surrounds the number once it's spoken. Swapping in a more senior engineer to run the same unstructured, inside-view process tends to produce a similarly biased number with more confidence attached to it, which is arguably worse.

The Anchor Gets Set Before the Analysis Starts

One specific mechanism deserves its own treatment because it's so common in software planning meetings and so rarely named: anchoring on the first number spoken.

Anchoring, another finding from the same research tradition that produced the planning fallacy, describes how an initial number (even one presented as a rough guess, even one the presenter explicitly flags as unreliable) becomes the reference point against which every subsequent adjustment gets measured, and those adjustments tend to be too small relative to the anchor. In an estimation meeting, the first number spoken is rarely the product of the most careful analysis in the room. It's often spoken early precisely because someone was asked a direct question and felt social pressure to answer quickly rather than to say "I need to think about this and get back to you."

Once that number exists, group dynamics reinforce it rather than interrogate it. A product manager who has just heard "probably six weeks" and needs to communicate a date to a customer or a board has an incentive to treat six weeks as a floor to defend, not a hypothesis to test. An engineer who privately thinks six weeks is optimistic faces a real social cost to saying so out loud in front of a room that has already, functionally, agreed; disagreeing now reads as obstruction rather than diligence, even though it's the more responsible position. Jørgensen's estimation research specifically documents that client- or manager-suggested effort or deadline expectations measurably pull expert estimates toward those expectations, independent of the technical facts of the task. This is not a claim that engineers are dishonest under pressure. It's a claim that the number in the room changes the number in people's heads, even for people trying to be accurate.

The practical consequence is that many software estimates are not really independent technical judgments. They are technical judgments partially anchored to whatever number was said earliest and loudest, adjusted at the margins by people who know better but who are adjusting from the wrong starting point. A process that wants a genuinely independent estimate has to structurally prevent this: for instance, by having each estimator write down a number privately before any number is spoken aloud, a practice closely related to the planning-poker technique used in some agile teams, precisely because it interrupts the anchor before it forms.

Why the Most-Cited "Projects Are Always Late" Statistic Is Worth Being Skeptical Of

Any article about software estimation eventually runs into the Standish Group's CHAOS Report, whose headline figures about project failure and cost overrun have been quoted in presentations, vendor pitches, and consulting decks for three decades. It's worth pausing on this specific statistic, because it's a useful lesson in exactly the kind of unverified claim this article is trying to avoid, and because the actual scholarly correction is more interesting than the myth.

Jørgensen and Kjetil Moløkken-Østvold published a detailed methodological critique of the CHAOS Report's famous claim that the average software project runs 189 percent over budget (The Rise and Fall of the Chaos Report Figures — Jørgensen & Moløkken-Østvold, published via Simula Research Laboratory). Their findings are worth summarizing plainly because they undercut a number that a great many engineering leaders still repeat as settled fact.

The Standish Group never disclosed its sampling methodology, reportedly stating that doing so would be "like giving away their business for free." The 189 percent figure itself was described inconsistently across the Standish Group's own publications: Jørgensen and Moløkken-Østvold found that of fifty documents citing the number, roughly half interpreted it as a cost overrun percentage and most of the rest interpreted it as cost expressed as a percentage of the original estimate, which is a mathematically different quantity. The original data-gathering process reportedly involved calling IT executives and asking them to share stories of project failure, which is a sampling method almost guaranteed to select for dramatic failures rather than a representative cross-section of projects. And tellingly, the Standish Group's own later surveys showed the reported overrun figure declining sharply over successive years (189 percent in 1994, 142 percent in 1996, 69 percent in 1998, 45 percent in 2000) eventually converging toward a range much closer to figures reported by more methodologically transparent academic studies from the 1980s and early 1990s, which put average software cost overruns in the neighborhood of 30 percent.

Two things follow from this. First, when a number sounds dramatic enough to have survived thirty years of repetition in slide decks, that durability is not evidence of accuracy: it's evidence that the number is memorable and rarely checked, which is exactly the situation this article's own production guidelines are built to avoid. Second, and more constructively, the more carefully sourced academic estimates (cost overruns in the range of roughly 30 percent, drawn from studies with disclosed methodology) describe something closer to what most engineering leaders actually experience: not catastrophic failure on the majority of projects, but a persistent, moderate, one-directional overrun that shows up often enough to be a pattern rather than bad luck. That pattern is consistent with a systematic bias, not with random variance, and it's consistent with the planning-fallacy mechanism described above. A moderate, one-directional, extremely common overrun is a more useful and more honest description of the problem than a dramatic, poorly sourced one, because it points toward a process fix instead of toward fatalism about how "everyone" runs over.

The Systematic Direction of the Error: What Gets Cut First

If estimation bias were random, the categories of work it hit would be random too. It isn't, and they aren't. Across the anecdotal experience of engineering leaders, the estimation research above, and the practical realities of how software gets built, a small set of categories reappear as the ones estimates habitually shortchange, and it is not a coincidence that they are also the categories most likely to be silently compressed once a deadline is under pressure.

The mechanism connecting these two facts is simple and worth stating directly: work that is hard to picture in advance is also, usually, work that is easy to skip once time runs short, because skipping it doesn't produce an obvious, immediate gap the way skipping a coded feature would. Nobody ships a feature that's visibly half-built; the code either exists or it doesn't, and stakeholders can see the difference. Nobody outside the engineering team can as easily see the difference between a feature that received two days of integration testing and one that received two hours, until the difference shows up in production. This asymmetry in visibility is why the same categories of work get underestimated at planning time and de-prioritized at crunch time: they're the same categories because they're invisible for the same reason at both points.

Work category Why it's invisible at estimation time Typical consequence when the schedule compresses
Integration testing across real dependencies The estimator pictures their own code, not how it behaves against another team's service, a third-party API's actual (not documented) behavior, or production-like data volumes Integration issues surface late, often after code freeze, forcing a choice between slipping the date or shipping with a known unresolved integration risk
Edge cases and account/state variations Edge cases are, by definition, not the case anyone is picturing when they imagine "the feature working" Edge-case handling gets triaged as "nice to have" and either cut from scope quietly or discovered as a defect after release
Regression testing after late changes Rarely estimated as its own line item; assumed to be covered by "testing" in general, which itself is often under-scoped Late-breaking changes near the deadline get shipped with reduced or no regression coverage, since the regression pass is the first thing cut to protect the release date
Code review and rework from review feedback Coding time is estimated as "time to write the code," with review treated as a formality rather than a phase that generates real rework Review gets rushed or skipped under time pressure, or rework from review eats into the time that was reserved for testing
Non-functional requirements (performance, accessibility, security review) These are rarely part of the feature description that prompted the estimate in the first place Deferred to "a later pass" that frequently never happens, or discovered as a production incident
Environment and tooling friction Assumed to be a solved problem because it usually is, until the one time a staging environment doesn't match production or a dependency version conflict eats a day Absorbed silently into whichever phase has the least political power to defend its own time, almost always testing
Requirements that were implied but never stated The person requesting the feature didn't think to mention them because they seemed obvious; the estimator didn't ask because the plan seemed complete Discovered during testing or acceptance review, at which point they read as "scope creep" even though they were part of the actual requirement from day one

Steve McConnell's Software Estimation: Demystifying the Black Art, a widely used reference in the software engineering management literature, makes a version of this same point from the practitioner side: many software estimates systematically exclude activities like code review, defect fixing, mentoring less experienced team members, and supporting other parts of the system, because those activities don't feel like "the work" when someone is asked to estimate a feature (Software Estimation: Demystifying the Black Art — Steve McConnell, O'Reilly / Microsoft Press). McConnell also draws a distinction that engineering leaders should treat as load-bearing: an estimate is a probabilistic forecast of effort under uncertainty, while a target is a business-driven date someone wants to hit, and a commitment is a promise the team makes about what they'll deliver by when. These are three different things that get collapsed into a single number in most planning meetings — someone asks for an estimate, receives a target dressed up as an estimate, and then treats it as a commitment, at which point the number has drifted through three different meanings without anyone noticing the shift.

The practical upshot is that "testing" as a line item in a project plan is frequently a euphemism for "whatever time remains after coding took as long as it took," not an independently reasoned estimate of how long it will actually take to find and fix the defects a feature of this complexity is likely to contain. That is the specific, mechanical reason testing time and delivery date collapse into the same failure. They were never actually two separate numbers to begin with: testing was defined from the start as the leftover.

Hypothetical example: the SaaS reporting feature that ate its own test window

Consider a hypothetical but realistic scenario at a mid-sized B2B SaaS company building a new usage-based billing report for enterprise customers.

Initial situation: The product team requests a self-serve report that lets enterprise customers see their usage broken down by team, project, and API endpoint, exportable as CSV. Engineering estimates three weeks, based on a mental model of building the aggregation query, the report UI, and the export function.

The hidden assumption: The estimate implicitly assumed the billing data was already clean and consistently structured across all customers, because that's how it looks in the primary database the engineer usually queries during development. Nobody explicitly said "we're assuming the data is clean": it simply wasn't part of anyone's mental picture of the work, because clean data is what engineers see every day in their own development environment.

The technical and organizational cause: In production, several of the company's oldest enterprise accounts have usage records from three different billing schema versions, the product of two prior migrations that were never fully backfilled. This is exactly the kind of fact that is invisible during the inside-view estimation exercise and only becomes visible once someone runs the report logic against real account data, which is a testing activity, not a coding activity.

The consequence: In week three, with the export function built and the release date already communicated to the customer success team, integration testing against production-like data reveals that the report silently produces incorrect totals for roughly a dozen of the company's largest accounts, precisely the accounts most likely to notice a billing discrepancy and escalate it.

The decision that needs to be made: Ship on the committed date with a known defect affecting large accounts, quietly exclude those accounts from the report's initial release, or slip the date to properly handle the schema variations. Each option has a real cost: reputational risk with the most important customers, an awkward partial rollout that erodes trust in the roadmap, or a delay that will be described in the next planning meeting as "testing took longer than expected."

The better approach: None of these three options was necessary if the estimate had separated "build the aggregation and export" from "validate against the actual variety of production billing data we support," as two distinct line items, with the second scoped by asking what data conditions genuinely exist across the customer base rather than by assuming the shape of the data the engineer sees locally. That's a five-minute conversation with whoever owns the billing data model, conducted before the estimate is finalized rather than discovered during testing. The three-week number was not wrong because the engineer was bad at estimating code. It was wrong because it never included a line item for a category of work (validating real data variety) that the estimator had no way to picture without asking a specific question first.

Why "Just Add a Buffer" Doesn't Fix a Systematically Biased Process

The most common organizational response to chronic lateness is to pad the estimate: add 20 percent, add a sprint, add a contingency line to the project plan. This intervention treats the problem as random noise around an otherwise accurate number, and it fails for a specific, structural reason: a systematic bias doesn't get corrected by a fixed multiplier, because the size and shape of the missing work varies by feature, by team, and by how many unknown unknowns a particular piece of work happens to contain. A flat buffer is a statistical fix applied to a non-statistical problem.

There's a second, more organizational reason buffers fail that shows up constantly in practice and is rarely named directly: a buffer that is visible to anyone besides the estimator gets negotiated away. This isn't cynicism about product managers or executives; it's a predictable consequence of incentives. If a team says a feature will take four weeks and privately believes three weeks is more likely with some risk of overrun, and if that buffer is visible as a labeled line in the plan ("1 week contingency"), it reads to anyone under their own delivery pressure as slack rather than as a hedge against a documented, research-backed bias. Slack gets reallocated. The one week gets pulled into an earlier deadline, or redirected to a different feature that's behind, and the team is back to an unbuffered estimate with an official record showing that they agreed to the tighter number.

This dynamic is closely related to what's sometimes discussed in project management as an application of Parkinson's Law: work expands to fill the time available, but the less-discussed corollary is that available time also contracts to match whatever pressure exists to reallocate it, particularly when the person applying the pressure doesn't bear the consequence of the contraction. The person deciding to pull a week of QA time forward to a different priority is rarely the person who will be paged when that missing week turns out to have mattered.

The consequence is that padding, applied naively, doesn't protect quality time. It protects whichever number is easiest to defend in a prioritization conversation, and an unlabeled, structural, itemized allocation for testing and integration is much harder to argue away than a line labeled "buffer," because "buffer" announces itself as optional in a way that "integration testing for the three external dependencies this feature touches" does not.

Hypothetical example: the fintech team whose buffer became someone else's deadline

A hypothetical fintech company is building a new reconciliation feature that matches incoming payment webhooks against internal ledger entries, a category of work with a well-earned reputation for edge cases (duplicate webhooks, out-of-order delivery, partial refunds, currency rounding).

Initial situation: The engineering lead, aware of this feature category's history, estimates five weeks of coding and explicitly adds a two-week buffer for reconciliation edge-case testing, presenting seven weeks total to the roadmap review.

The hidden assumption: The lead assumed that labeling the buffer explicitly, as "reconciliation edge-case hardening," would protect it, because it was named rather than hidden inside a padded overall number.

The technical and organizational cause: Three weeks into the project, a different, unrelated initiative falls behind, and the VP of Product asks whether the reconciliation feature's date can move up to free engineering capacity for the other initiative. In the resulting conversation, the two-week buffer is the only clearly discretionary-looking line in the plan: the five weeks of coding is not something anyone believes can be compressed without cutting scope, but the buffer, by its own label, announces itself as time that isn't tied to a specific deliverable yet.

The consequence: The buffer is cut to one week under the reasoning that "we can always add more testing later if we find something," and the release ships on the revised six-week date. Reconciliation defects involving out-of-order webhook delivery surface in production three weeks later, requiring an emergency patch and a manual reconciliation of affected accounts, a more expensive and higher-risk activity than the testing time that was cut would have been.

The decision that needs to be made, and the better approach: The fix is not to stop communicating contingency to stakeholders: hiding risk from the business is its own failure mode. The fix is to make the itemized nature of the work, not a generic buffer label, the thing being defended. Instead of "five weeks coding, two weeks buffer," the plan should show "five weeks coding; two weeks for testing duplicate delivery, out-of-order delivery, partial refund, and currency-rounding scenarios specifically, based on known reconciliation failure modes from [a named prior incident or a documented category of risk]." A line item tied to specific, named risk is much harder to characterize as slack in a reprioritization conversation than a line item called "buffer," even though the underlying time allocation might be identical. The problem in this scenario wasn't that the lead failed to protect quality time. It's that the protection mechanism chosen (a generically labeled contingency) was structurally the easiest thing in the plan to argue away.

Reference-Class Forecasting: Using the Team's Own History Instead of Its Imagination

If the inside view is the mechanism that produces the bias, the outside view (reference-class forecasting) is the most evidence-backed correction available, and it adapts to software teams more directly than it might first appear.

The method, as Flyvbjerg has applied it to infrastructure megaprojects and as transportation and public-works agencies have since adopted it for cost estimation, has three steps: identify a reference class of past projects genuinely comparable to the one being planned, establish the actual distribution of outcomes for that reference class (not the plan, the actual result), and use that distribution (not a fresh inside-view analysis of the new project) as the primary basis for the forecast, adjusted only for specific, evidenced differences between the new project and the reference class (Reference class forecasting: promises, problems, and a research agenda moving forward; From Nobel Prize to Project Management: Getting Risks Right — PMI).

Applied to a software team, this looks less like a formal government methodology and more like a discipline the team can build with data it likely already has, if it starts tracking it deliberately.

Building a reference class in practice:

  1. Define comparable categories of past work, not a single undifferentiated bucket of "all our past tickets." A reference class of "features that touch three or more services" behaves very differently from "internal admin tooling changes," and collapsing them together produces a forecast that's wrong for both.
  2. Record actual outcomes, not just planned ones, for a rolling set of recent work in each category: planned effort, actual effort, and (critically for this article's argument) actual effort specifically for testing, integration, and defect-fixing phases, tracked separately from coding effort.
  3. Compute the ratio, not just the difference, between planned and actual for each category. A multiplicative uplift factor (actual was typically 1.4x the original estimate for this category) generalizes better across different-sized features than an additive one (add two weeks), because the missing work scales with the size and complexity of the feature, not with a fixed constant.
  4. Apply the reference-class ratio as the starting forecast, then adjust only for a documented, specific reason the current feature differs from the reference class, not for optimism about the current team's focus or a general feeling that "this one should be simpler."
  5. Feed the actual outcome of the current feature back into the reference class once it ships, so the data set improves rather than staying fixed at whatever sample happened to be available when the practice started.

Illustrative worked example (hypothetical figures, not a benchmark)

To make this concrete: suppose an engineering team building customer-facing SaaS integrations reviews its last twelve features that involved a new third-party API integration, a reasonably specific and comparable reference class. It finds that the original estimate-to-actual ratio for total effort averaged 1.6, and when the team breaks that gap down by phase, it finds the coding phase came in close to estimate (a ratio of about 1.1), while the combined integration-testing-and-defect-fixing phase came in at a ratio of roughly 2.3 relative to what had been budgeted for it, in several cases because "integration testing" hadn't been given its own line item at all and had been implicitly folded into a generic "testing" allowance sized as an afterthought.

For a new, thirteenth integration feature that the team's inside-view analysis estimates at four weeks of coding and one week of testing, the reference class suggests a materially different shape: roughly 4.4 weeks of coding (a modest adjustment) and closer to 2.3 weeks of integration testing and defect-fixing, not one week: nearly seven weeks total rather than the five the inside view proposed, with the gap concentrated almost entirely in the category the original estimate treated as an afterthought. These figures are illustrative, constructed to demonstrate the calculation method, not a benchmark drawn from real industry data, and any team applying this method should build its ratios from its own tracked history rather than borrowing a number from an example in an article.

The output of this exercise is not a magic corrected number. It's a forecast built from what the team's own history says actually happens to plans like this one, instead of a forecast built from what the team can currently picture happening. That's the entire mechanism by which reference-class forecasting outperforms inside-view estimation, and it requires no new statistical sophistication: only the discipline of tracking actuals by category and being willing to let history override optimism.

Teams too new to have twelve comparable past features can still apply the principle at a smaller scale: even three or four tracked examples of "how much longer did integration testing take than we planned for this kind of work" is a better anchor than an unadjusted inside-view guess, and it's a practice that compounds: the reference class only gets more useful the longer a team maintains it.

Making Testing and Integration Visible as Explicit Line Items

The research points to a specific structural fix, distinct from adding a generic buffer: stop letting "testing" and "integration" hide inside a single number called "development," and require them to be estimated as their own line items, by the people who will actually do that work, using the same rigor applied to coding time.

This sounds like a small procedural change and behaves like a significant one, because the act of itemizing forces exactly the kind of specific question ("what data conditions does this feature need to be validated against?") that the SaaS billing example above shows was skipped precisely because nobody was required to answer it before the estimate was finalized.

A practical line-item template for feature estimation:

Line item What it should explicitly account for Who should estimate it
Implementation Writing the code for the primary, expected-case behavior The engineer(s) building the feature
Integration validation Verifying the feature's behavior against real or realistic instances of every external dependency it touches (other services, third-party APIs, production-shaped data) The engineer(s) building the feature, informed by whoever owns the dependency being integrated
Edge-case and negative-path testing Account states, input variations, and failure conditions outside the primary expected case QA or the engineer, working from an explicit list of known edge-case categories for this type of feature, not from memory alone
Regression validation Confirming that existing, adjacent functionality still behaves correctly after the change, especially for late-breaking modifications QA, or the engineer if there is no dedicated QA function
Defect remediation Time to fix what the above phases find, genuinely unknown in advance, but estimable as a range based on the team's own reference-class history for similar work Whoever will be doing the fixing, informed by historical defect rates for comparable features
Non-functional validation (if applicable) Performance under realistic load, accessibility, security-relevant behavior, where the feature genuinely implicates any of these The relevant specialist, or the engineer working from a documented checklist for this feature category

This template does not need to be applied with equal weight to every ticket: a two-line copy change does not need six line items, and forcing that level of ceremony onto trivial work is its own failure mode, producing process fatigue that makes teams resent and eventually abandon the practice. It's the right tool for features with genuine integration surface, meaningful edge-case exposure, or anything touching money, security, or data integrity: precisely the features where the SaaS billing and fintech reconciliation examples above show the inside view's blind spot does the most damage.

The itemized template also solves the political problem the earlier buffer example ran into. "Two weeks buffer" is a number anyone can propose cutting without needing to understand what's inside it. "1.5 weeks for integration validation against three external services, this feature's known-highest-risk category based on our last six integrations" is a number that requires a specific, informed counter-argument to cut: someone has to argue that the integration risk doesn't apply this time, which is a much higher bar than arguing that a generic contingency line looks like slack.

A checklist for introducing line-item estimation without adding process overhead

  1. Apply it selectively. Reserve the full line-item breakdown for features with real integration surface, edge-case exposure, or risk to money, security, auth, or data integrity. Estimate low-risk, well-understood work the way the team already does.
  2. Name the specific risk, not a generic category. "Integration testing" as a line item is better than nothing; "validate against the three billing schema versions still present in production" is far more defensible and far more accurate.
  3. Separate the estimator from the anchor. Where practical, have the person estimating integration and testing time do so before hearing anyone's overall target date, to reduce the anchoring effect described earlier.
  4. Require a one-line justification for any adjustment that shrinks a line item after the fact. Not to create bureaucracy, but to make the moment of cutting testing time a visible decision rather than a silent default.
  5. Record what actually happened, per line item, once the feature ships. This is the raw material the reference-class method above depends on, and skipping this step is the single most common reason estimation practices don't improve over time even when teams intend them to.
  6. Revisit the categories periodically. A team's known risk categories should evolve: a category that caused problems in the past (a particular flaky third-party API, a specific legacy data format) deserves an explicit line until the underlying risk is actually resolved, not indefinitely, but not forgotten after one bad quarter either.

Tracking Estimation Accuracy as a Metric in Its Own Right

Most engineering organizations track delivery metrics of some kind: velocity, cycle time, deployment frequency, change failure rate. Few deliberately track how accurate their estimates actually were, as a distinct, ongoing measurement, separate from whether a particular sprint felt busy or a particular release felt chaotic. This is a gap worth closing, because everything in this article up to this point depends on a team being able to say, with actual data, how its estimates have historically compared to outcomes; the entire reference-class method is impossible without it.

Software delivery predictability is a theme that shows up in the broader DevOps research literature as well. Google's DORA research program, which produces the annual Accelerate State of DevOps Report, has documented that unstable organizational priorities measurably reduce delivery performance and increase burnout, and that improvement is most reliable when teams take an experimental approach: establishing a baseline, forming a specific hypothesis about what to change, and measuring whether the change actually moved the number, rather than adopting practices because they sound generically good (Accelerate State of DevOps Report 2024 — DORA / Google Cloud). Estimation accuracy tracking is a natural extension of that same discipline, applied to a metric DORA's four core measures (deployment frequency, lead time for changes, change failure rate, and time to restore service) don't directly capture: not how fast or how stable delivery is, but how honest the team's forecasts of its own delivery have been.

A useful estimation-accuracy practice does not need to be elaborate. At minimum, for each estimated unit of work above a reasonable size threshold, record the original estimate, the actual outcome, and (where the team has adopted line-item estimation) the actual outcome broken down by category. Reviewed quarterly, this produces a genuinely useful signal: is the team's estimation bias getting smaller, staying flat, or (worth catching early) getting worse as the codebase and its dependencies grow more complex?

It's worth being explicit about which versions of this metric are useful and which are actively misleading, because a metric introduced carelessly can create the same perverse incentives that engineering leaders already worry about with velocity or story-point tracking.

Approach to tracking estimation accuracy Why it helps Why it can backfire if misused
Tracking planned-vs-actual ratio by work category (coding, integration, testing) over a rolling window, reviewed at the team level Surfaces exactly where the bias concentrates, feeds the reference-class method, improves over time without needing anyone to be "right" on any single estimate None significant if kept at the team level and framed as a forecasting tool, not a performance measure
Publishing individual engineers' estimation accuracy as an individual performance metric Sounds like accountability Produces padded, defensive estimates and discourages anyone from taking on genuinely uncertain work, since uncertain work is exactly what produces a "bad" accuracy score through no fault of the estimator
Treating a single estimate's miss as evidence the process failed Feels intuitive Individual estimates will always vary; the useful signal is in the aggregate direction and magnitude across many estimates, not any one data point
Using aggregate historical ratios to set forecast ranges for future work, openly shared with stakeholders as ranges rather than single dates Sets more honest expectations, reduces the anchoring dynamic described earlier because a range is harder to treat as a single fixed target Requires stakeholders willing to plan around a range instead of demanding a single number, which is a genuine organizational change, not just a spreadsheet change
Measuring "percentage of releases where testing time was reduced from what was originally scoped" Directly measures the specific failure mode this article describes Needs a consistent definition of "originally scoped" testing time to be meaningful, which is only possible once line-item estimation is in place

The last row deserves attention on its own, because it's arguably the single most diagnostic metric available to a leader trying to find out whether their organization has this specific problem. If a team can look back over its last dozen releases and find that testing time was cut from its original allocation in most of them, that is direct, unambiguous evidence of the mechanism this article describes, independent of whether any individual release felt like a crisis at the time.

A Framework for Protecting Quality Time When the Deadline Is Already Fixed

Everything above addresses how to estimate better going forward. It does nothing for the team that is, right now, holding a fixed deadline and a testing budget that was never adequate to begin with, because the estimate that produced that deadline already happened, weeks or months ago, and can't be redone. This situation is not exceptional. Given everything described above about how estimation bias works, it should be treated as close to the default condition any team should expect to face at some point in most delivery cycles, and having a deliberate framework ready for it is more useful than treating it as a surprise every time.

The mistake most teams make in this moment is letting the compression happen implicitly: testing time simply shrinks as a side effect of the deadline holding firm and the coding work taking as long as the coding work takes, with nobody making an actual decision about what gets cut. The alternative is to make the trade-off explicit and structured, even under real time pressure, using a consistent order of what can be compressed and what genuinely cannot.

The Compression Order: a decision sequence for a fixed deadline and an inadequate remaining testing budget

When a team realizes, partway through a delivery cycle, that the remaining time before the deadline will not accommodate the testing and hardening work the feature actually needs, the following sequence should be worked through explicitly, in order, with each step documented as a decision rather than allowed to happen by default:

  1. Confirm the deadline is actually fixed, and by whom, for what reason. A surprising number of "fixed" deadlines turn out to be internally assumed rather than externally committed: nobody has actually told a customer or regulator a specific date, and the fixed feeling comes from an internal roadmap slide rather than an external obligation. If the deadline is genuinely externally committed (a contractual date, a regulatory deadline, a announced customer-facing date), treat it as fixed and move to step two. If it isn't, the first and cheapest option is renegotiating the date itself, which is almost always less costly than either of the remaining options.
  2. Cut scope before cutting testing. Reducing what ships is a visible, deliberate, reversible decision that stakeholders can evaluate and disagree with in the open. Reducing testing coverage is usually invisible to anyone outside engineering until it produces a defect, which makes it a much worse place to absorb pressure precisely because the cost is deferred and hidden rather than immediate and negotiated. If any part of the feature's scope is genuinely deferrable (a secondary use case, a less-common account type, a nice-to-have export format) deferring it and testing a smaller surface thoroughly is close to always the better trade than testing a larger surface superficially.
  3. If scope truly cannot be cut, rank remaining testing work by blast radius, not by how long it takes. Not all testing work protects against equally severe outcomes. Testing that protects against data loss, financial miscalculation, security exposure, or irreversible customer-facing mistakes should never be the first thing cut, regardless of how much time it would save, because the cost of a failure in these categories is usually far higher than the cost of a further schedule slip. Testing that protects against a cosmetic UI inconsistency on a rarely used screen is a legitimate candidate for reduction under real time pressure.
  4. Make any reduction in testing coverage an explicit, recorded decision with a named owner, not a default. Whoever decides that regression testing for a lower-risk area will be reduced should be identifiable, and the decision should be written down, even briefly. This does two things: it forces the decision to actually be made consciously rather than happening as an unexamined default, and it creates the accountability that makes people more careful about making it lightly.
  5. Communicate the trade-off honestly to whoever owns the business risk, not just to engineering leadership. A product leader or executive who is told "we're reducing test coverage on the bulk CSV export feature to hit this date, and the residual risk is X" can make an informed call about whether that trade-off is acceptable. A product leader who isn't told anything and only learns about the reduced coverage when a defect surfaces in production has effectively had that decision made for them without their knowledge, which erodes trust in a way that a difficult but transparent conversation does not.
  6. After the release, feed the actual outcome back into the reference class described earlier. If the reduced testing area held up fine, that's useful data about which categories of risk this team can safely compress under pressure. If it didn't, that's equally useful data about which categories cannot be compressed again without a different plan.

This sequence does not eliminate the underlying problem: an inadequate testing window is still an inadequate testing window, and no amount of prioritization turns two days into five. What it does is convert an invisible, default erosion of quality time into a visible, deliberate, accountable set of decisions, made by the people who should be making them, with the actual risk understood by the people who bear it. That is a materially different outcome than the SaaS billing example earlier in this article, where nobody decided to skip validating the billing data; it just happened, because nobody made the decision explicit until it was too late to make a different one.

Hypothetical example: the e-commerce checkout redesign under a fixed launch date

Initial situation: An e-commerce company is redesigning its checkout flow ahead of a peak shopping season, with a launch date fixed by the marketing calendar and already communicated externally through a planned promotional campaign.

The hidden assumption: The original nine-week estimate assumed the new checkout would integrate against the existing payment gateway with no meaningful changes to that integration, since the redesign was scoped as a UI and flow change, not a payments change.

The technical and organizational cause: During integration testing in week seven, the team discovers that the new checkout's asynchronous form validation introduces a timing change in how payment authorization requests are sent, occasionally causing the payment gateway to receive a duplicate authorization request under specific network-latency conditions, a defect category the original estimate had no line item for, because nobody expected a UI redesign to touch payment-request timing at all.

The consequence: With two weeks remaining before a fixed, externally communicated launch date, the team faces exactly the situation this framework addresses: an inadequate remaining window to fully test and remediate a defect category that touches money, discovered too late to simply extend the original estimate.

The decision that needs to be made: Working through the Compression Order, the team first checks whether the date can move: it can't, given the promotional campaign already announced to customers. Scope reduction is considered next: the team identifies that a secondary "save payment method for next time" feature, bundled into the same redesign, can be deferred without affecting the core checkout flow, freeing several days. This alone isn't enough. Applying the blast-radius ranking, the team explicitly decides to fully test the duplicate-authorization scenario across all supported payment methods (the highest-risk category, since it involves customers potentially being charged twice) while reducing visual regression testing on less-used checkout layout variants (a cosmetic risk category, not a financial one) to a lighter spot-check rather than full coverage.

The better approach: This decision, made explicitly and documented with the VP of Product and the head of Customer Support informed of the reduced visual-regression coverage and the specific residual risk it carries, is a fundamentally different outcome than the same reduction happening silently because nobody had time to get to it. The launch proceeds on schedule, the duplicate-authorization defect is caught and fixed before release because it was correctly prioritized as the highest-blast-radius risk, and a minor visual inconsistency on one rarely used checkout variant is caught and patched two days after launch with essentially no customer impact, a trade-off the business consciously accepted rather than one that happened to it.

Applying This Differently by Company Stage

The mechanisms described in this article apply everywhere, but the practical response looks different depending on how mature and how large an engineering organization is.

Startups and early-stage teams usually don't have twelve past features to build a reference class from, and shouldn't try to force a heavyweight process onto a five-person team. The realistic version of this practice at early stage is qualitative: after each release, spend ten minutes explicitly asking "what took longer than we thought, and was it coding or was it testing and integration?" and keeping even an informal record of the answer. The goal at this stage is building the habit of separating those two categories in the team's own thinking, so that by the time the company has enough history to formalize reference-class forecasting, the underlying discipline of tracking the distinction is already normal.

Scale-ups with established engineering teams are usually the best-positioned to get immediate value from the full practice described in this article: enough historical ticket data to build genuine reference classes, enough process maturity to introduce line-item estimation without it feeling like bureaucratic overhead, and often the most acute version of the underlying pain, since scale-ups are frequently the stage where a growing customer base makes production defects newly expensive at exactly the moment engineering process is still informal. This is the stage where introducing an estimation-accuracy metric, reviewed quarterly by engineering leadership, tends to produce the clearest return, because the organization is large enough for patterns to be statistically meaningful but still small enough to change process quickly.

Enterprises typically already have more formal estimation processes, sometimes overly formal ones (story points, planning poker, capacity models) but the same underlying inside-view bias can hide inside a very rigorous-looking process just as easily as an informal one, because the bias operates at the level of what the estimator can picture, not at the level of process ceremony. A large enterprise's most valuable move is often auditing whether its existing formal estimation process actually itemizes testing and integration as distinct categories with their own tracked accuracy, or whether it has simply formalized the same inside-view, coding-centric estimate with more governance around it. Enterprises are also the organizations most likely to have genuine reference-class data buried across many teams and past projects that nobody has aggregated; often the single highest-leverage fix available is simply building the reporting to make that existing data visible and usable, rather than introducing an entirely new estimation methodology.

Questions Engineering and Product Leaders Should Bring Back to Their Teams

Before the next estimation meeting, a leader who wants to interrupt this pattern rather than repeat it can usefully ask:

  • When we say a number for this feature, is testing and integration time included in that number, or is it something we'll figure out once the coding is done?
  • If we look at our last several comparable features, how did the actual testing and integration time compare to what we originally planned for it, and do we actually have that data, or are we guessing?
  • Whose number was said first in this meeting, and did everyone else's estimate get anchored to it before they had a chance to think independently?
  • If this deadline turns out to be too tight once we're partway through, do we have an agreed way to decide what gets cut, or will it just happen by default to whatever has the least visible advocate in the room?
  • Are we tracking estimation accuracy as its own signal at all, separate from whether a given sprint felt hard?

Software Estimation Accuracy Is a Process Property, Not a Talent Property

The organizations that improve at this don't do so by finding better estimators. They do so by changing what an estimate is required to include, by separating the number a technical person produces from the target a business stakeholder wants, by tracking their own history closely enough to use it instead of imagination, and by treating the moment a deadline gets tight as a decision to be made deliberately rather than a default to be absorbed silently by whichever phase of work is least able to defend itself.

The uncomfortable part of this argument, worth stating directly rather than softening, is that "we shipped late" and "we shipped buggy" are usually not two separate risks a team is managing. They're two possible outcomes of the same underestimated number, and a team under enough pressure will often get both at once (a shorter-than-hoped slip combined with a riskier-than-intended release) because the compressed testing time doesn't fully compensate for the missed original date, either. The real choice most teams are making, whether they realize it or not, is not "on time versus late." It's "which visible consequence do we prefer for the same underlying estimation failure." Naming that choice explicitly, before it happens by default, is the actual improvement available here, more available, and more within a team's control, than getting better at predicting the future.

The question worth taking back to any engineering organization is not "why are we always late." It's "what would our estimates look like if testing and integration were priced with the same rigor as the code, and are we willing to find out."

Teams that recognize this pattern in their own release history and don't yet have the tracked data to break it are often better served by an outside perspective that has seen the same failure mode across many different codebases and can help build the testing and integration discipline into the estimate itself, rather than discovering it during the crunch. QAtronic's software testing services work with engineering teams to make integration testing, edge-case validation, and regression coverage explicit, estimable parts of a delivery plan rather than an afterthought absorbed under deadline pressure, and where a team simply needs additional QA capacity to protect a testing window that a fixed deadline would otherwise compress, QAtronic's outstaffing and staff augmentation options can add that capacity without requiring a permanent headcount decision.

Frequently Asked Questions

Is the planning fallacy the same thing as being bad at estimating? No. The planning fallacy describes a systematic, direction-specific bias that affects experienced and inexperienced estimators alike, because it stems from how people mentally construct a plan (the inside view), not from a lack of technical skill. Improving individual technical judgment does not, by itself, correct a bias rooted in what the estimation process asks people to picture.

Why is testing specifically the category that gets underestimated, rather than some other part of the work? Testing and integration work is, by nature, the process of discovering problems nobody has pictured yet. The inside view used to build most estimates can only account for what the estimator can currently imagine, so the category of work whose entire purpose is finding the unimagined is structurally the hardest category to size accurately in advance.

Does adding a time buffer to an estimate solve this problem? Not reliably. A buffer applied as a flat percentage doesn't match the actual, variable size of the missing work, and a buffer that's visible to stakeholders as a discretionary line item is frequently the first thing removed under later scheduling pressure, since it's easier to argue is optional than an itemized, named allocation for specific integration or edge-case work.

What is reference-class forecasting, in practical terms, for a software team? It means estimating a new piece of work by starting from how similar past work actually turned out, rather than by reasoning only from the details of the current plan. In practice, this requires tracking actual outcomes (not just original estimates) for past features by category, particularly separating testing and integration time from coding time, and using the resulting ratios as the starting point for new estimates.

How do we track estimation accuracy without it becoming a way to punish individual engineers? Track it in aggregate, at the team or work-category level, reviewed periodically as a forecasting input, not as an individual performance measure tied to any single estimate. Publishing individual accuracy scores tends to produce defensive, padded estimates and discourages engineers from taking on genuinely uncertain work, which defeats the purpose of the metric.

What should we do if we discover mid-project that our original testing time was never enough, and the deadline can't move? Work through the trade-off explicitly rather than letting it happen by default: confirm whether the deadline is genuinely fixed, look for scope that can be deferred before touching testing coverage, prioritize remaining testing by the severity of what could go wrong rather than by how long each test would take, and make any reduction in coverage a documented decision that the business stakeholder who owns the risk is actually informed about.

Is this only a problem for large, complex projects, or does it affect small features too? The mechanism applies at any scale, but the consequence scales with the feature's blast radius. A small internal tool with an underestimated testing phase produces a minor inconvenience; a small feature that touches payments, authentication, or customer data with the same underestimated testing phase can produce a serious incident despite its small size, which is why blast radius, not size, should drive how much rigor a given estimate gets.

How is this different from just getting better at agile story-point estimation? Story points are typically a relative-sizing technique for coding effort and are vulnerable to the same inside-view bias described in this article if testing and integration aren't explicitly included in what a point is meant to represent. The fix described here is compatible with any estimation technique a team already uses; it changes what gets included in the estimate and how accuracy gets tracked over time, not the specific numerical method used to express the estimate.

Recent posts

October 2, 2026
FHIR Interoperability Testing: Certified, Not Connected
October 2, 2026
Definition of Done Erosion: Why Standards Quietly Slip
October 2, 2026
Kubernetes Admission Control Testing: A Field Guide