Kubernetes Resource Limits Need Their Own Test Plan
Share this post

The OOMKill Nobody Load-Tested For: Why Kubernetes Resource Limits Need Their Own Test Plan

A payments API team once spent a sprint shipping what looked like a routine improvement: an in-memory cache in front of a slow currency-conversion lookup, added to shave latency off a hot endpoint. Every test passed. Unit tests passed. The integration suite passed. QA ran the full regression pass in staging and signed off. The load test that had been run when the service first went into production, eighteen months earlier, had never been rerun, because nothing about the observable behavior of the endpoint had changed enough to trigger a retest. The service deployed cleanly. Two days later, under the next business-hours traffic ramp — hypothetically, since this scenario is illustrative rather than a real QAtronic engagement — the pod started restarting every eleven minutes. Latency spiked right before each restart, then recovered, then degraded again. The on-call engineer saw CrashLoopBackOff, checked application logs, found nothing resembling an error, and eventually ran kubectl describe pod to see the line that explained everything: Reason: OOMKilled, Exit Code: 137.

Nothing in that story involves a defect in the traditional sense. The code did exactly what it was written to do. It cached the right values, returned the right conversions, and passed every test written against its behavior. The problem was that the cache had no eviction policy tuned to the container's actual memory ceiling, and nobody had a test whose job was to check whether the service, under real concurrency and a realistic cache fill pattern, still fit inside the memory limit set in the deployment manifest a year and a half earlier. That test didn't fail. It didn't exist.

This is a narrower and more specific problem than "we need better observability" or "we should do more chaos testing." It is a gap in test ownership: for most teams running production workloads on Kubernetes, there is no test — automated or manual, in CI or in staging — whose explicit job is to verify that a service's real memory and CPU behavior under realistic sustained load still fits the resource limits Kubernetes has been told to enforce, and that this remains true after the code inside the container changes. Functional tests validate correctness. Most performance tests run once, early, often in an environment with generous or entirely unset limits, and are treated as a launch gate rather than a recurring check. Neither one is positioned to catch a service that works correctly but has quietly outgrown the box it was told to live in.

Why This Keeps Slipping Past Every Existing Gate

The interesting part of an OOMKill incident is rarely the incident itself. Kubernetes does exactly what it is designed to do: a container exceeds its declared memory limit, the kernel's control-group memory controller detects it, and the container's process is terminated by the kernel's out-of-memory killer. Kubernetes marks the container OOMKilled, the pod's restart policy typically brings it back, and if the underlying cause hasn't been addressed, the cycle repeats. This behavior is documented directly in Kubernetes' own resource-management and memory-assignment references, which describe requests as a scheduling signal and limits as a hard ceiling enforced by the kubelet through the container runtime's cgroup configuration, with an exceeded memory limit resulting in termination and the pod exiting with code 137 (see Kubernetes: Assign Memory Resources to Containers and Pods and Kubernetes: Resource Management for Pods and Containers).

The part worth examining is not the kill itself but the sequence of test gates the service passed on its way to production, none of which were positioned to catch it.

Start with functional and integration testing. These suites run against correctness: does the endpoint return the right data, does the workflow complete, does the edge case get handled. They typically run in ephemeral environments, on a handful of requests, for a few seconds each. A memory leak that accumulates over hours of sustained traffic, or a cache that grows unbounded only once real-world key cardinality is reached, is structurally invisible to a test that spins up a container, sends ten requests, and tears it down. The test isn't wrong. It's simply answering a different question than "does this fit in the box."

Next, staging environments. Many teams configure staging with generous memory limits, or none at all, specifically to avoid staging becoming a source of false-positive failures unrelated to the feature being tested. That's a reasonable choice for a QA environment focused on functional coverage. But it means staging actively hides the one failure mode this article is about: a service that behaves identically in staging and production except for the one variable staging was configured to ignore.

Then the load test that happened once, early. Most teams that run performance testing at all run it before launch, or during a major redesign, then treat the results as durable. QAtronic's own writing on this class of problem — the gap between a single-point validation and how a system behaves months later — applies directly here: a load test run against last year's code, before three dependency upgrades, a caching layer, and a larger default batch size, tells you almost nothing about this year's memory footprint under the same traffic pattern. The load test wasn't wrong when it ran. It just wasn't rerun.

Finally, production monitoring. Dashboards typically track CPU and memory utilization as a percentage or an absolute value sampled every fifteen to sixty seconds. A memory spike that triggers a cgroup OOM kill can happen inside a window shorter than the metrics collection interval — the container can allocate past its limit and be killed before the monitoring system ever records the peak. Google Cloud's own guidance for diagnosing out-of-memory events on GKE makes exactly this point, warning engineers not to rely solely on memory utilization metrics from monitoring tools because collection intervals routinely miss the rapid spikes that trigger a kill (see Google Cloud: Troubleshoot OOM events in GKE). By the time a dashboard shows a plausible memory graph, the smoking gun has already been cleaned up by the restart.

Put those four gates together and the failure mode becomes clear: correctness tests don't look at memory footprint, staging doesn't enforce realistic limits, load tests go stale, and production monitoring often can't see the moment of failure clearly enough to diagnose it without extra digging. A service can pass through all four checkpoints, be technically correct, and still fail in a way that only becomes visible as a pattern of restarts under real production concurrency.

What Kubernetes Actually Enforces, and Why the Enforcement Is Stricter Than Most Teams Assume

Resource limits are frequently treated as an infrastructure knob — something the platform team sets once during onboarding, based on a rule of thumb or whatever the previous service used. Understanding why that approach breaks down requires being specific about what Kubernetes is actually doing under the hood, because the enforcement mechanism is less forgiving than the phrase "resource limit" tends to suggest.

A pod specification separates requests from limits, and the two serve different purposes. The request is what the scheduler uses to decide which node has room for the pod — it is a reservation, not a ceiling. The limit is what the kubelet enforces at runtime through the container's control group (cgroup), and it is a hard ceiling: the container is not allowed to exceed it, and Kubernetes' own documentation is explicit that this enforcement happens at the kernel level, independent of anything the application does or is aware of (Kubernetes: Resource Management for Pods and Containers).

CPU and memory are enforced differently, and the difference matters for how a service fails. A CPU limit is enforced through throttling: if a container tries to use more CPU than its limit allows, the kernel's CFS bandwidth controller restricts its CPU time in that scheduling period. The container doesn't crash — it slows down, sometimes badly, in ways that show up as latency spikes rather than restarts. A memory limit is enforced categorically differently, because memory cannot be throttled the way CPU can — a process either has the memory it's trying to allocate or it doesn't. When a container's memory usage reaches its memory.max cgroup value and the kernel cannot reclaim enough memory to bring usage back under that ceiling, the kernel's OOM killer selects a process within that cgroup and terminates it. The Linux kernel's own cgroup v2 documentation states this plainly: memory.max is "the main mechanism to limit memory usage of a cgroup. If a cgroup's memory usage reaches this limit and can't be reduced, the OOM killer is invoked in the cgroup" (Linux Kernel Documentation: Control Group v2, Memory). This is why a CPU-bound problem tends to look like degraded performance, while a memory-bound problem tends to look like a service that is fine and then abruptly is not.

It's also worth distinguishing two failure modes that get lumped together under "OOM" in casual conversation, because they point to different root causes and require different fixes. GKE's troubleshooting documentation for OOM events separates them clearly:

  • Container-level (cgroup) OOM — a single container exceeds its own declared memory limit. The kernel's memory controller kills processes inside that container's cgroup specifically. Kubernetes records this as OOMKilled and, depending on the pod's restart policy, brings the container back up. This is the scenario this article focuses on — it's isolated to one workload and is directly addressable by fixing the code, the limit, or both.
  • Node-level (system) OOM — the entire node runs out of available memory across every process on it, not just one container's cgroup. The kernel's global OOM killer selects a victim process system-wide, which can be a different, unrelated workload than the one that actually caused the memory pressure. This is messier to diagnose because the pod that gets killed isn't necessarily the pod that caused the problem, and it often indicates that node-level resource accounting, bin-packing, or eviction thresholds need attention rather than any single service's code.

(See Google Cloud: Troubleshoot OOM events in GKE for the full breakdown, including log query patterns for each type.)

There's a second layer of enforcement worth understanding, because it changes how a service behaves under pressure even before it gets OOMKilled: Kubernetes' Quality of Service (QoS) classification. A pod is automatically assigned one of three QoS classes based on how its requests and limits are configured:

QoS class How it's assigned Behavior under node memory pressure
Guaranteed Every container specifies both CPU and memory requests and limits, with requests equal to limits Evicted last; effectively shielded unless the node itself is critically starved
Burstable At least one container specifies a request, but requests are below limits (or limits are unset for one resource) Evicted based on how far actual usage exceeds the request, after all BestEffort pods
BestEffort No requests or limits specified at all Evicted first under any memory pressure

(Source: Kubernetes: Configure Quality of Service for Pods.)

This matters for the article's central argument because it means the "resource limits" conversation isn't only about the ceiling a single container hits — it's also about how that container's declared shape affects its standing relative to every other workload on the node the moment memory gets tight. A Burstable service that requests 256Mi but is regularly observed using 900Mi isn't just risking its own OOMKill; it's also a disproportionately attractive eviction target the moment a neighboring pod puts pressure on the node, because the kubelet's eviction logic weighs actual usage against the declared request, not against what the service "really needs."

There is a third enforcement layer worth understanding, because it explains why a service can degrade for other reasons before an individual container ever hits its own memory limit: kubelet node-pressure eviction. This is a separate mechanism from the kernel's cgroup OOM killer, and it operates proactively rather than reactively. The kubelet continuously monitors node-level signals — most importantly memory.available — and when those signals cross a configured threshold, it begins evicting pods to reclaim resources before the node runs out entirely. Kubernetes distinguishes two threshold types here: soft thresholds, which allow a grace period so pods can shut down cleanly, and hard thresholds, which trigger immediate eviction with no grace period at all, on the assumption that waiting would risk starving the whole node (Kubernetes: Node-pressure Eviction). Crucially, this is a kubelet-level decision informed by QoS class and relative resource usage across every pod on the node — it is not the same event as a single container's cgroup hitting its own memory.max and being killed by the kernel. A service can be evicted by node-pressure eviction without ever individually exceeding its own declared limit, simply because it happened to be a lower-priority, higher-usage-relative-to-request tenant on a node that got tight on memory for an unrelated reason. This is one more reason a single container's memory graph, viewed in isolation, doesn't tell the whole story — the node's aggregate pressure and the mix of QoS classes sharing it matter just as much as any one service's own ceiling.

CPU limits deserve a similar clarification, because teams sometimes assume CPU and memory misconfiguration produce comparable symptoms when they don't. A container that exceeds its CPU limit is throttled by the kernel's CFS bandwidth controller for the remainder of that scheduling period — its execution is paused, not terminated. The practical effect is elevated latency and, in request-handling services, a queue of work that backs up during the throttled window, rather than a restart. This matters for the diagnostic side of this article's argument: a service showing latency spikes and timeouts without any pod restarts is very unlikely to be hitting its memory limit at all, and investigating it as a memory problem wastes time that should go toward checking CPU throttling metrics instead. The two failure modes look different enough in practice that conflating them slows down root-causing either one.

A minimal example makes the requests/limits relationship concrete:

yaml
resources:
  requests:
    memory: "256Mi"
    cpu: "250m"
  limits:
    memory: "512Mi"
    cpu: "500m"

This is a Burstable pod. It's guaranteed 256Mi and can burst up to 512Mi before the kubelet's cgroup enforcement kills it. Whether 512Mi is a safe ceiling for this particular service, under this particular traffic pattern, after this particular code change, is precisely the question that no functional test, and usually no re-run load test, is set up to answer.

How a Fully Passing Test Suite Coexists With a Growing Memory Footprint

It helps to separate the actual causes of a service outgrowing its limit from the reasons those causes go undetected, because the two are often confused. The causes are ordinary engineering changes, not bugs:

A new caching layer. Caching is one of the most common ways a service's steady-state memory footprint increases, because the entire design goal is to hold data in memory that would otherwise be fetched again. An in-process LRU cache with no explicit size bound, or a size bound expressed in entry count rather than bytes, can grow well past what its designer estimated once real production key cardinality is reached — which is often far higher than whatever sample of keys was used to sanity-check it locally.

A larger default batch size. Batch processing services — nightly reconciliation jobs, webhook fan-out workers, export generators — often have a batch size or page size that gets bumped up for throughput reasons: fewer round trips, better amortized overhead, faster completion. Each row in a larger batch holds its own place in memory until the batch is processed and released. A batch size increase that looks like a two-line configuration change can multiply peak memory linearly, and that multiplication is invisible to a functional test that validates correctness on a batch of ten records rather than ten thousand.

A dependency upgrade. Library and runtime upgrades routinely change default buffer sizes, connection pool defaults, internal caching behavior, or garbage collection tuning, usually for good reasons — better throughput, safer defaults, new features. An HTTP client library that changes its default connection pool size, or a JSON parser that switches to a faster but more memory-hungry buffering strategy, can shift a service's baseline memory usage without a single line of application code changing. Dependency upgrades are exactly the kind of change that sails through code review as "just a version bump" and through CI as "all tests green," because the tests were never asserting anything about memory footprint in the first place.

A slow leak that only manifests under sustained load. Some memory growth isn't a step change but a gradual accumulation — an event listener that isn't unsubscribed, a connection that isn't returned to its pool under a specific error path, a metrics library accumulating label cardinality over time. These only become visible after hours of continuous operation, which is longer than almost any pre-deploy test run and shorter than the interval between most teams' periodic load tests.

None of these four causes is a coding mistake in the sense QA traditionally hunts for. They're reasonable engineering decisions, made independently, by people who had no reason to think of them as resource-limit changes — because from the application developer's point of view, they aren't. They're caching decisions, throughput decisions, and dependency decisions. The resource-limit consequence is a side effect nobody was assigned to watch for.

It's worth noting what these four causes have in common structurally, because it points toward how to catch the next one without waiting for a full taxonomy of every possible cause. Each one increases the amount of data a service holds in memory per unit of work, per connection, or per unit of time — more cached entries, more records per batch, larger default buffers, or memory that simply isn't released. None of them changes what the service computes or returns; all of them change how much memory that computation requires along the way. That's a useful mental filter for code review: a change that doesn't alter what a service does, but plausibly alters how much it holds in memory while doing it, is exactly the category this article is about, regardless of whether it falls neatly into "caching," "batching," "dependency," or some other label.

Hypothetical Scenario One: The Cache That Ate the Headroom (Fintech)

Situation. A fintech company runs a currency-conversion microservice behind its core payments API. The service has run stably for eighteen months on a Burstable pod configuration: 256Mi request, 512Mi limit. A platform engineer added a small in-memory cache in front of an external rate-lookup call to cut p95 latency, using a popular caching library with a default of "unbounded unless configured otherwise."

Hidden assumption. The engineer assumed the number of distinct currency pairs the service would ever cache was small — a few hundred, at most — because that's what the company's supported currency list looked like at the time. Nobody documented or tested that assumption; it was implicit in not setting an explicit cache size bound.

Technical/organizational cause. Six months later, the product team added support for a partner integration that passed through dozens of regional sub-currencies and promotional pricing tiers as distinct cache keys, multiplying real key cardinality by roughly twentyfold. No one working on that partner integration knew a cache existed three services upstream, let alone that it had no eviction policy.

Consequence. The service's steady-state memory usage climbed gradually over about three weeks as the cache filled with the new key space, crossing the 512Mi limit during peak partner traffic and triggering repeated OOMKilled restarts specifically during the partner's business hours — a pattern that looked, at first glance, like a partner-side issue rather than an internal one.

Decision to be made. Whether to simply raise the memory limit (a quick fix that defers the same problem to the next cardinality increase) or to bound the cache explicitly and add a test that asserts memory behavior under realistic key cardinality.

Better approach. Bound the cache by memory footprint rather than entry count, add an explicit eviction policy, and add a load test scenario that simulates realistic key cardinality — including the partner's key space — as a required check before any change that touches the caching layer, not just before major releases. Just as important as the technical fix is the organizational one: the partner-integration team needs a lightweight way to know that a caching layer exists upstream of the fields they're adding, which in practice usually means documenting cache-sensitive fields in the service's interface contract rather than expecting every consuming team to trace the call graph themselves.

Hypothetical Scenario Two-and-a-Half: The Leak That Only Showed Up After Hours (Enterprise SaaS Support Platform)

Situation. An enterprise SaaS company runs a long-lived WebSocket connection service that keeps support agents' dashboards updated in real time. The service had been stable in staging through every pre-release test, each of which ran for a few minutes at most before tearing the environment down.

Hidden assumption. The team assumed that because the service passed every staging test and every short synthetic load test, it was safe to treat as production-ready — without anyone questioning whether a few minutes of test duration was long enough to represent a service meant to run continuously for weeks between deploys.

Technical/organizational cause. A recently added per-connection metrics collector registered a new label combination on every WebSocket connect event but never cleaned up label state on disconnect, a common category of mistake with metrics libraries that use unbounded label cardinality. Over a few minutes this was invisible. Over several days of continuous operation with thousands of connect/disconnect cycles, the metrics library's internal state grew steadily, consuming memory that was never released.

Consequence. The service ran fine for roughly four to five days after each deploy before the accumulated metrics-label memory pushed it past its limit, producing a restart that reset the leaking state and bought another four or five days before the next one — a slow, low-frequency pattern that was easy to miss against the noise of normal operational restarts until someone plotted restart timestamps against deploy timestamps and noticed they didn't correlate with deploys at all, but with elapsed uptime.

Decision to be made. Whether to accept periodic restarts as an acceptable operating pattern (since the service degraded gracefully via reconnect logic on the client side) or to treat the elapsed-time-correlated pattern as a leak worth fixing at the source.

Better approach. Recognize that a restart pattern correlated with uptime rather than deploys or traffic is close to a diagnostic signature for a slow leak, and extend at least one load test scenario to run for a duration long enough — hours, not minutes — to surface this class of problem before it reaches a service with real-time client connections where every restart briefly drops active sessions.

Hypothetical Scenario Two: The Batch Size That Multiplied Under Load (E-commerce SaaS)

Situation. An e-commerce analytics SaaS company runs a nightly order-reconciliation job as a Kubernetes CronJob, processing merchant transaction batches. A backend engineer increased the batch size from 500 to 5,000 records per iteration to cut the job's total runtime, after confirming correctness against a test merchant account with a few hundred orders.

Hidden assumption. The engineer assumed per-record memory overhead was small and constant, so a 10x batch size increase would translate to a modest, proportional memory increase — reasonable for the record shape the test account had, which excluded any line-item detail or refund history.

Technical/organizational cause. In production, the largest merchant accounts carried orders with deeply nested line-item and refund-history objects, each holding significantly more memory per record than the lightweight test fixtures. A 10x batch size increase against that record shape produced a peak memory footprint roughly 25 times the previous job's peak, because per-record cost wasn't actually constant across the merchant population.

Consequence. The reconciliation job ran successfully against every merchant except the handful of largest accounts, whose pods were OOMKilled partway through processing — silently, because the CronJob's failure alerting only fired on job-level failure after several retries, and the partial reconciliation state for those merchants was inconsistent until someone noticed the discrepancy days later.

Decision to be made. Whether to treat this as a one-off batch-size rollback, or to build a repeatable performance check that specifically exercises the largest realistic record shapes in the merchant population, not just the most common ones.

Better approach. Test batch-processing changes against a record-shape distribution that includes the heaviest realistic cases, not only typical ones, and assert peak memory usage explicitly as part of that test — with the assertion tied to the job's actual configured resource limit, not to an assumption about "should be fine." This is also a case where the earlier QAtronic discussion of jobs that silently fail rather than loudly erroring is directly relevant — see Scheduled Job Reliability: Cron Jobs Fail Silently for the broader pattern of batch and scheduled work failing in ways that don't trip standard alerting.

Hypothetical Scenario Three: The Dependency Upgrade With a New Default (Enterprise SaaS)

Situation. An enterprise SaaS company's platform team performs a routine quarterly dependency upgrade cycle across its microservices, including bumping a widely used HTTP client library to its latest major version in a customer-facing notification service. The upgrade's changelog describes performance improvements and bug fixes; nothing in it explicitly flags a memory-relevant change.

Hidden assumption. The team assumed that because the service's automated test suite passed and the changelog didn't mention breaking changes, the upgrade was functionally and operationally equivalent to the previous version — a reasonable assumption for correctness, but not one that was ever actually validated for resource footprint.

Technical/organizational cause. The new library version shipped with a larger default connection pool and more aggressive internal response buffering, intended to improve throughput for typical workloads. The notification service, which maintains many long-lived outbound connections to customer webhook endpoints, saw its steady-state memory baseline rise by a meaningful margin under normal operating conditions — not because of a leak, but because the new default configuration simply held more in memory per connection than before.

Consequence. The service's baseline memory usage sat closer to its configured limit immediately after the upgrade, without failing yet. It took a subsequent traffic spike — a customer running a bulk notification campaign — to push it over the edge, weeks after the dependency upgrade had shipped, making the upgrade far from the first place anyone thought to look.

Consequence for root-cause analysis. Because the OOMKill happened weeks after the actual triggering change, the incident review initially focused on the traffic spike as the cause, and only a deeper investigation traced the shifted baseline back to the dependency upgrade's changed defaults.

Decision to be made. Whether dependency upgrades should be treated as resource-neutral by default (the status quo at this company) or whether upgrades affecting core runtime libraries — HTTP clients, serialization libraries, database drivers — should trigger a mandatory memory-footprint comparison against the previous version before merging.

Better approach. Treat dependency upgrades to core runtime libraries as a category of change that requires a short, targeted load test comparing memory footprint before and after, specifically because changelogs describe intended functional changes and rarely describe incidental resource-footprint changes. This doesn't require the full performance test suite — a focused before/after comparison under representative load is enough to catch a shifted baseline before it becomes an incident.

These three scenarios share a structure worth naming explicitly: in each one, the change that caused the eventual OOMKill was reviewed, tested for correctness, and merged by people acting reasonably, and the resource-limit consequence became visible only under a load pattern, data shape, or elapsed time that none of the existing test gates were built to exercise.

The Ownership Gap: Whose Test Plan Is This, Actually?

Ask five people at a mid-sized engineering organization who owns verifying that a service's memory usage still fits its Kubernetes limit after a code change, and it's common to get five different, half-right answers. The platform or DevOps team will often say it's the application team's responsibility, since they wrote the code that changed. The application team will often say resource limits are a platform concern, since the platform team set the original values and owns the Kubernetes configuration. QA will often say performance testing is out of scope for a typical feature ticket, and that the last dedicated load test happened at launch. Everyone is describing a real boundary of their role. None of them is wrong about their own scope. The result is that a specific, concrete kind of testing has no home.

Function What they typically own What falls outside their usual scope
Application/feature team Correctness of the feature, unit and integration tests, code review Whether the change alters steady-state or peak memory footprint under realistic load
Platform/DevOps/SRE team Cluster configuration, initial resource limit values, autoscaling policy, node capacity Re-validating that limits still fit the workload after every material application change
QA/performance testing function Functional regression, and often a one-time load test near launch Recurring memory-footprint verification tied to specific code changes (caching, batching, dependency upgrades)
Production monitoring/on-call Detecting and responding to active incidents Root-causing which change introduced a shifted memory baseline weeks earlier

The gap isn't a staffing problem in the sense of needing more people. It's a definitional problem: none of these functions has been assigned a specific, recurring trigger — "when X kind of change happens, someone runs Y kind of test against Z kind of load, and checks the result against the declared limit." Without that trigger, the responsibility diffuses to "whoever notices during an incident," which is exactly what happened in all three hypothetical scenarios above.

Closing this gap doesn't require creating a new team. It requires making an explicit decision about which existing function owns a specific, narrow test: verifying declared resource limits against real memory and CPU behavior under realistic sustained load, triggered by a defined set of change types rather than by calendar cadence alone. In most organizations that already have a performance testing function or work with an external QA and DevOps consulting services partner, this fits most naturally as a shared responsibility — platform engineering owns the limit values and the enforcement configuration, while the performance testing function owns the recurring verification that those values still hold, with both sides agreeing on what triggers a re-check.

Writing that responsibility down explicitly, even briefly, is what actually closes the gap — not the org chart itself. A useful exercise is assigning one word to each function for this one specific activity: who is responsible for running the test, who is accountable for the limit values being correct, who needs to be consulted before a triggering change ships, and who simply needs to be informed when a test fails. In most of the incidents described in this article, that assignment had never been made explicitly for this narrow activity — resource limits were "owned" by platform engineering in the general sense of who wrote the YAML, but no one was responsible, specifically, for re-verifying them against new code. Making that assignment doesn't require new tooling or a new hire. It requires a short conversation and a line in a runbook or team charter that says, plainly, which function reruns this test and what triggers it.

What Each Existing Test Gate Actually Catches — and What It Doesn't

It's worth being precise about this, because the instinct after an OOMKill incident is often to add more of something teams already have — more monitoring, more load testing "eventually," a bigger limit as a buffer. The table below is meant to make the actual coverage gap visible rather than assumed.

Test gate What it verifies Typical environment Does it catch a resource-limit regression?
Unit tests Function-level correctness Local / CI, no real container limits No — doesn't run inside a resource-constrained container at all
Integration tests Cross-service correctness, short-lived requests CI or ephemeral namespace, often generous or unset limits Rarely — short duration and small data volume miss growth patterns
Staging functional QA Feature behaves correctly end-to-end Staging, frequently configured with loose limits to avoid unrelated failures No — staging is often deliberately configured to hide this class of failure
One-time launch load test Throughput and latency at expected scale, at time of launch Pre-production, matched to code as it existed at that time No, after the first material change — the test goes stale the moment the code it validated changes
Production monitoring dashboards Aggregate resource utilization trends Production Partially — often misses short spikes faster than the metrics collection interval, and rarely attributes a shift to a specific prior change
Recurring resource-limit test plan (the gap this article addresses) Actual memory/CPU behavior under realistic sustained load, checked against declared limits, triggered by defined change types A dedicated test environment configured to mirror production limits Yes — by design, this is the only gate whose explicit job is this question

Reading the table left to right, the pattern is that every existing gate is doing its assigned job correctly. The absence of the sixth row is what allows the incident, not a defect in rows one through five.

This reframing matters for how a postmortem gets written after an incident like this. It's tempting to write "QA missed this" or "our load testing wasn't thorough enough" as the root cause, but that framing blames a gate for not doing a job it was never assigned. QA's functional suite did exactly what functional suites are supposed to do. The launch load test did exactly what a launch load test is supposed to do, at the time it was run. A more accurate root cause, and a more actionable one, is that no gate was ever assigned the specific, recurring job of re-verifying resource-limit fit after the kind of change that actually caused the incident. Writing the postmortem that way points toward the actual fix — creating and assigning the missing gate — rather than toward vague instructions to "test more thoroughly" that rarely survive past the next sprint planning meeting.

Building a Resource Limit Test Plan That Actually Gets Triggered

A test plan that only exists in principle, without a defined trigger for when it runs, tends to decay into the same one-time-at-launch pattern this article is arguing against. The following framework is built around two things: what to test, and — more importantly — what specifically triggers the test, since the trigger is usually the part that's missing.

Define the trigger conditions first. A recurring resource-limit test should not depend on someone remembering to run it. It should be triggered by specific, identifiable change types:

  1. A change to caching behavior — a new cache, a change to eviction policy, or a change to what gets cached.
  2. A change to batch size, page size, or any configuration governing how much data is processed or held in memory per operation.
  3. An upgrade to a core runtime dependency — HTTP clients, database drivers, serialization libraries, web frameworks, or the language runtime itself.
  4. A change to concurrency configuration — worker counts, connection pool sizes, thread pool sizes.
  5. A meaningful increase in expected traffic or data volume — a new large customer, a new integration, a marketing campaign with a known traffic profile.
  6. Elapsed time since the last verification — a backstop for changes that don't fall cleanly into the categories above, set to a cadence appropriate to how frequently the service changes (a fast-moving service might warrant quarterly checks; a stable one might warrant semiannual checks).

Treat a canary or progressive rollout as a complement, not a substitute, for pre-deploy testing. A canary deployment — releasing a change to a small percentage of production traffic before a full rollout — can genuinely help catch a resource-limit regression before it affects every user, and many teams already have this mechanism in place for exactly this kind of risk reduction. But it's worth being honest about its limits here: a canary running for a short window will catch a fast memory spike but is much less likely to catch a slow leak that only manifests after hours, and by definition a canary is a small slice of real users experiencing degraded behavior while the signal accumulates. It's a good second line of defense, not a reason to skip pre-deploy testing for changes that match the trigger conditions above.

Design the load pattern to match reality, not convenience. The test needs sustained load, not a burst — long enough to reveal gradual memory growth, not just an initial spike. For most services this means running the test for a duration measured in tens of minutes to a few hours at realistic concurrency, rather than the few-minute smoke tests common in CI. It also means using realistic data shapes — the heaviest reasonable record sizes and key cardinality the service will actually see in production, not the smallest fixture that makes the test fast.

Run it against the actual declared limit, not a generous test-environment limit. This is the detail most likely to be skipped, and it's the one that makes the biggest difference. If the test environment gives the container 2Gi "to avoid noisy test failures" while production enforces 512Mi, the test will pass regardless of whether the service fits in production. The test environment's resource configuration needs to mirror the production manifest for this specific test to mean anything.

Assert against the limit explicitly, not just observe. A load test that produces a memory graph someone has to eyeball is much weaker than one that fails automatically when memory usage approaches the declared limit. A simple approach many performance testing tools support is a threshold assertion tied to a safety margin below the hard limit — for example, using a tool like k6 with a custom check against metrics scraped from the container, or asserting against Prometheus query results in a CI step:

javascript
// Example k6 threshold pattern — fails the test run if
// container memory (scraped via a custom trend metric)
// exceeds 85% of the declared limit during sustained load.
export const options = {
  thresholds: {
    'container_memory_bytes': ['p(95)<430080000'], // 85% of 512Mi limit
  },
};

The exact tooling matters less than the principle: the test should produce a pass/fail result tied to the declared limit, with a safety margin, rather than a chart that requires a human to notice a trend.

Include an intentional over-limit run as a control. It's useful to occasionally verify that the test itself is capable of catching a real regression — for example, temporarily reverting an eviction policy or bumping a batch size in a test branch to confirm the test fails as expected. A resource-limit test that has never actually failed in its history is worth treating with some suspicion; it may be testing the wrong thing.

Decide where this test lives in the pipeline. A resource-limit test that requires sustained load for tens of minutes generally doesn't belong in the same fast-feedback CI stage as unit tests — it belongs in a dedicated performance stage, gated on the trigger conditions above rather than run on every commit. For teams already running a nightly or pre-release performance suite, the most practical integration path is usually adding resource-limit assertions to that existing suite rather than standing up an entirely separate pipeline. The goal is to make the check cheap enough to run whenever a trigger condition fires, not to make every engineer wait tens of minutes for it on every pull request.

Keep the test environment's node characteristics honest, not just its resource limits. Matching the container's declared memory limit is necessary but not sufficient — if the test cluster's nodes have dramatically more headroom than production nodes, node-pressure eviction behavior (described earlier) won't manifest the same way, and a service that would be an early eviction target in production under real node contention may look perfectly healthy in a spacious test cluster. Where practical, testing on nodes sized and packed similarly to production gives a more honest signal than testing on an oversized dedicated node that happens to be convenient.

A Practical Checklist: Does a Change Need a Resource-Limit Retest?

Use this before merging a change, as a quick gate rather than a full load-test cycle for every pull request:

  • Does this change add, remove, or reconfigure any in-memory cache or in-process data structure that grows with usage?
  • Does this change alter batch size, page size, buffer size, or any "how much data at once" configuration?
  • Does this change upgrade a core runtime dependency (HTTP client, database driver, serializer, framework, language runtime)?
  • Does this change alter worker count, connection pool size, or thread pool configuration?
  • Does this change target a service expected to see a meaningful increase in traffic, data volume, or customer scale in the near term?
  • Has it been longer than the service's defined retest interval since the last resource-limit verification?

If any box is checked, the change should trigger the recurring resource-limit test before it reaches production, not after the first incident makes the need obvious.

Diagnosing an OOMKill That's Already Happening

Some readers arrived here mid-incident rather than in a planning cycle. The diagnostic sequence below is built around distinguishing the failure modes described earlier, since the fix differs depending on which one is actually happening.

Step one: confirm it's actually an OOMKill, and which kind. Run kubectl describe pod <pod-name> and look at the Last State section for Reason: OOMKilled and Exit Code: 137 — this confirms a container-level cgroup kill rather than an application crash, a liveness probe failure, or a node-level system OOM event (which would typically show up differently, often as the pod simply disappearing from a node that itself reports memory pressure). If the evidence points to a node-level event rather than a single container hitting its own limit, the investigation should shift toward node-level bin-packing and capacity rather than a single service's code — GKE's OOM troubleshooting guide provides log query patterns to distinguish the two cases directly (Google Cloud: Troubleshoot OOM events in GKE).

Step two: recover the evidence from before the kill. The application logs from the crashed container are gone from the live pod's log stream by the time anyone investigates, but Kubernetes retains the previous container's logs for exactly this situation:

bash
kubectl logs --previous <pod-name>

This is the single most useful command in this diagnostic sequence, and it's also the one most commonly skipped in favor of staring at the live pod's current, uninformative logs (see Kubernetes: Debug Running Pods).

Step three: distinguish a leak from a spike from a step change. These require different fixes.

  • A gradual leak shows memory climbing steadily over hours before the kill, visible in dashboards despite their sampling limitations, and typically points toward unclosed resources, unbounded accumulation structures, or a specific error path that skips cleanup.
  • A sudden spike shows memory jumping sharply right before the kill, often too fast for standard dashboard sampling to capture clearly, and typically points toward a specific request pattern — a large payload, a pathological query, a burst of concurrent requests hitting a shared in-memory structure at once.
  • A step change in baseline shows memory settling at a new, higher steady state after a specific deploy, rather than continuously climbing, and points toward a configuration or dependency change that altered default behavior — exactly the pattern in the dependency-upgrade scenario earlier in this article.

Step four: correlate against recent changes, not just recent traffic. Because the triggering change and the resulting incident are often separated by days or weeks (as in the dependency-upgrade scenario), it's worth explicitly checking the deploy history for the service against categories from the checklist above — caching changes, batch size changes, dependency upgrades, concurrency changes — rather than assuming the most recent deploy is automatically the cause.

Step five: check whether the application's own memory ceiling matches the container's. For runtimes with their own internal heap or memory limit configuration — JVM heap size flags, Node.js --max-old-space-size, Go's GOMEMLIMIT — a mismatch between the application-level limit and the container's cgroup limit is a common, easy-to-miss cause. If the application believes it has more room than the container actually allows, the container's cgroup enforcement wins every time, and the application never gets a chance to garbage-collect proactively before the kernel intervenes.

Step six: check whether horizontal autoscaling has been masking the problem. A service with horizontal pod autoscaling configured on CPU or request-rate metrics can absorb a growing per-pod memory footprint for a surprisingly long time by simply running more replicas — each individual pod still gets killed and restarted periodically, but the aggregate capacity keeps up with traffic well enough that nobody notices a pattern until either the autoscaler hits its maximum replica count or the restart frequency itself starts affecting tail latency. If restart counts have been climbing gradually over weeks while overall service health looked fine, this is a common explanation, and it means the underlying memory regression may have been present, and growing, for far longer than the incident timeline suggests.

Step seven: check sidecar and init container overhead. In a service mesh or logging-sidecar architecture, the application container isn't the only consumer of the pod's resource budget in some configurations, and sidecar memory usage — particularly under a traffic spike that also increases mesh proxy buffering — can be an underappreciated contributor to a pod-level memory ceiling being reached.

Setting Limits That Can Actually Be Tested Against

None of the diagnostic or test-plan work above matters much if the limits themselves are set in a way that makes them untestable in practice — either so tight that ordinary variance trips them constantly, or so loose that they never trip regardless of what the code does, which defeats their purpose as a real safety mechanism.

Avoid copy-pasted values across services with different workloads. It's common for a new service's resource limits to be copied from an existing service's manifest as a starting point — a reasonable way to avoid starting from zero, but risky as a permanent value if nobody revisits it once the new service's actual behavior diverges from the one it was copied from. A batch worker and a request-handling API have fundamentally different memory profiles even if they happen to share a starting template.

Use Guaranteed QoS deliberately for workloads where predictability matters more than density. Setting requests equal to limits (Guaranteed QoS) trades cluster bin-packing efficiency for eviction priority and predictable behavior — appropriate for a payments-critical service, less necessary for a low-priority background worker where occasional eviction under pressure is an acceptable trade-off. This is a real architectural decision, not a default to apply uniformly.

Treat the Vertical Pod Autoscaler as a feedback signal, not a substitute for testing. Kubernetes' Vertical Pod Autoscaler (VPA) can observe actual historical usage and recommend — or in some configurations automatically apply — updated requests and limits based on what a workload has actually consumed (Kubernetes: Vertical Pod Autoscaling). This is genuinely useful as an ongoing feedback loop, but it's reactive by construction: it recommends based on what already happened, which means it will always lag behind a code change until that change has already run in production long enough to generate a usage history — potentially including the incident itself. VPA and a proactive pre-deploy resource-limit test answer different questions and are complementary, not interchangeable; relying on VPA alone still means the first sign of trouble is a production incident.

Build in deliberate headroom, and document why. A limit set exactly at observed peak usage leaves no room for the normal variance between test conditions and production reality. A common, reasonable approach is to set the limit at a documented margin above the highest realistic sustained usage observed in testing — enough to absorb normal variance without being so loose that a genuine regression goes unnoticed for months.

Distinguish requests from limits deliberately, rather than defaulting one from the other. It's common to see manifests where the request and limit are set identically out of convenience, or where the request is left unset entirely and only a limit is specified. Neither is automatically wrong, but each has consequences: setting requests equal to limits produces Guaranteed QoS, trading flexibility for predictability, as covered earlier; leaving requests unset (with only a limit specified) still produces a functioning pod, but it removes the scheduler's ability to reason about how much of that resource the pod actually needs on a typical day, which can lead to poor bin-packing decisions across the cluster even when no single pod ever gets OOMKilled. Setting both values deliberately, with a documented rationale for the gap between them, is worth the extra few minutes it takes during initial configuration.

Revisit limits as part of change management, not as a one-off configuration task. The single biggest structural fix implied by everything in this article is treating resource limits the way a mature engineering organization treats API contracts or database schemas: as something with an owner, a documented rationale, and a defined process for when it needs to change — not as a value set once during onboarding and left alone until something breaks. A limit with no documented reasoning behind it is, in practice, a limit nobody actually owns, which is precisely the ownership gap this article opened with.

When a Dedicated Resource-Limit Test Plan Is Overkill

Not every service justifies the full framework above, and it's worth being honest about where the cost outweighs the benefit, since over-applying a rigorous process is its own kind of engineering waste.

A short-lived internal tool used by a handful of employees, with low blast radius if it restarts, generally doesn't need a recurring resource-limit test plan — a generous limit and standard monitoring are proportionate. A batch job with a completion deadline measured in hours, where an occasional restart-and-retry doesn't threaten an SLA, can often tolerate a looser limit and a simpler alert-on-repeated-failure approach rather than a full pre-deploy test cycle. A very early-stage startup with a handful of services and a small team may reasonably decide that the engineering time this framework requires is better spent elsewhere until the service portfolio and traffic scale justify it — the risk calculus for a two-person engineering team pre-product-market-fit is different from a scale-up serving paying enterprise customers with contractual uptime commitments.

The distinguishing question isn't service size or team size directly — it's blast radius and recovery cost. A customer-facing service on the critical path of a paying transaction, a service whose OOMKill loop degrades a shared dependency for other services, or a service whose restart causes data inconsistency rather than a clean retry, all justify the investment regardless of company stage. A background job that can safely retry from where it left off, serving no external SLA, usually doesn't — at least not until its failure starts actually costing something measurable.

What This Looks Like at Different Stages of Scale

The framework above doesn't need to be applied identically everywhere, and pretending otherwise is a common way well-intentioned process ends up ignored. What's proportionate differs meaningfully between a small team shipping fast, a scale-up with paying enterprise customers, and a large organization with dozens of services and multiple platform teams.

Stage Realistic resource-limit testing posture Common failure pattern to watch for
Early-stage startup (roughly one to two dozen services, small team) Lightweight: a documented reason for each limit, generous but not unbounded headroom, and a manual retest before any change explicitly flagged as caching- or batch-related. Full recurring automated test plans are often not worth the setup cost yet. Limits copied from a template with no one able to explain why the number is what it is; the first real signal is often the first incident.
Scale-up (meaningful paying customer base, dedicated platform or SRE function emerging) The checklist-triggered retest described in this article, applied at minimum to customer-facing and revenue-critical services, with staging configured to mirror production limits for those services specifically. Resource-limit testing exists for the two or three services that already had an incident, but hasn't been generalized to the rest of the fleet — coverage is reactive rather than systematic.
Enterprise (many services, multiple teams, formal SRE or platform organization) Fully triggered, automated resource-limit verification integrated into CI/CD for services above a defined criticality tier, with VPA or equivalent observability feeding a continuous baseline and periodic manual review for services below that tier. The process exists on paper across the whole organization, but enforcement varies by team, and a newly acquired or newly onboarded service can slip through without ever being brought into the standard the rest of the fleet follows.

The pattern across all three stages is the same even though the tooling differs: the organizations that avoid repeat incidents are the ones that wrote down, in some form, which changes trigger a retest and who runs it — not necessarily the ones with the most sophisticated tooling. A startup with a one-line rule in its deploy checklist ("bumping a batch size or adding a cache requires a manual memory check before merge") can be better protected than an enterprise with an elaborate performance testing pipeline that nobody has connected to a trigger condition.

Questions for Engineering Leaders to Bring Back to Their Teams

A short set of questions surfaces most of this gap quickly, without requiring a full audit:

  • When was the last time resource limits were actually retested against real load for our three or four most critical services — not just observed in a dashboard, but deliberately tested?
  • Which of our recent dependency upgrades, caching changes, or batch-size changes went through any process that considered their memory-footprint impact?
  • Does our staging environment enforce the same resource limits as production, or does it quietly allow more room?
  • If a service starts restarting under memory pressure next quarter, do we know who owns diagnosing it, and do they know to check kubectl logs --previous before the evidence rotates away?
  • Are our resource limits documented with a reason, or are they inherited values nobody currently on the team chose deliberately?

None of these require a large investment to answer. Most require an honest half-hour conversation across the platform and application teams — and the answers usually reveal which of the four causes described earlier in this article is most likely to bite next.

Weighing the Cost of Testing Against the Cost of Not Testing

It's worth being explicit about the trade-off being proposed here, rather than assuming the case for a recurring resource-limit test plan is self-evident. Building and maintaining sustained-load tests tied to specific trigger conditions takes engineering time — writing realistic test data generators, maintaining a test environment configured to mirror production limits, and reviewing results when a test does fail. That cost is real and shouldn't be waved away.

The table below is an illustrative framing, not a benchmark drawn from any real study or QAtronic engagement — the specific figures are placeholders meant to make the trade-off concrete, and any real organization should substitute its own numbers rather than treat these as representative.

Factor Illustrative one-time cost of building the test plan for a critical service Illustrative recurring cost per triggered retest
Engineering time A few days to define trigger conditions, build a realistic load profile, and wire up assertions A small fraction of a day per retest, once the harness exists
Infrastructure A test environment configured to mirror production limits (may already exist as part of staging) Minimal incremental cost if the environment is reused
What it avoids — An unplanned incident: on-call time, customer-facing degradation, and the investigation time needed to trace a regression back to a change that may be weeks old

Framed this way, the comparison isn't really "testing costs time" versus "not testing is free." It's a few days of upfront, plannable engineering work against an unplanned, unbounded cost that shows up later, at a worse time, and is harder to diagnose because the evidence has often rotated out of the logs by the time anyone starts looking. That asymmetry — plannable cost now versus unplanned cost later, with the added tax of investigation time — is usually the more persuasive argument internally than an abstract appeal to reliability.

The Distinction Worth Keeping

The instinct after an OOMKill incident is often to treat it as an infrastructure problem — bump the limit, add a node pool, move on. Sometimes that's the right immediate fix. But it treats the symptom as the disease. The actual gap is that a piece of behavior — how much memory a service actually uses under realistic sustained load — changed, and nothing in the pipeline was assigned to notice before production did. Raising the limit without adding the test simply moves the same unmonitored variable to a higher number, where it will eventually be crossed again by the next cache, the next batch size increase, or the next dependency upgrade.

The useful distinction isn't between teams that have OOMKill incidents and teams that don't — sustained-load memory behavior is genuinely hard to predict perfectly, and even well-tested services will occasionally surprise their owners. The useful distinction is between teams where an OOMKill incident triggers a new automated test tied to a specific trigger condition, and teams where it triggers a one-time limit increase and a Slack thread that gets forgotten by the next quarter. The first group's resource limits get more accurate over time. The second group's stay exactly as fragile as they were, waiting for the next change that happens to cross the same invisible line.

QAtronic works with engineering teams to build exactly this kind of missing test coverage — designing recurring performance and resource-limit verification that's triggered by real code and infrastructure changes rather than run once at launch and forgotten, as part of broader DevOps consulting services and performance testing engagements. For a team that already knows its last load test predates its last three dependency upgrades, that's usually the more useful starting conversation than a general infrastructure review — see QAtronic's approach to performance testing for how this fits into a broader test strategy.

Frequently Asked Questions

What does OOMKilled actually mean in Kubernetes? It means a container exceeded the memory limit declared in its resource configuration, and the Linux kernel's out-of-memory killer terminated the process inside that container's control group as a result. Kubernetes reports this as Reason: OOMKilled with Exit Code: 137 when you run kubectl describe pod.

Why didn't our load test catch this before launch? Most load tests run once, near launch, against the code as it existed at that time. If the service's caching behavior, batch sizes, dependency versions, or concurrency configuration change afterward — which is normal, ongoing engineering work — the original load test no longer reflects the current code, and nothing automatically reruns it.

Should we just increase the memory limit to stop the restarts? That can be a reasonable immediate mitigation, but by itself it doesn't address why the service's memory usage grew in the first place, and it doesn't add any mechanism to catch the next regression. A durable fix pairs the limit adjustment with a documented reason and, ideally, a recurring test that verifies the new limit still holds after future changes.

Does the Vertical Pod Autoscaler solve this problem for us? It helps, but it's a reactive feedback mechanism — it recommends limit adjustments based on observed historical usage, which means it only "learns" about a regression after that regression has already run in production. It's a useful complement to proactive pre-deploy testing, not a replacement for it.

Is this the same thing as chaos engineering? Not quite. Chaos engineering typically injects deliberate failures — killing pods, introducing latency, simulating network partitions — to test how a system recovers. A resource-limit test plan is narrower: it verifies that a specific, unmodified service's actual resource consumption under realistic load still fits its declared configuration. The two are complementary but answer different questions.

How often should we retest resource limits? Ideally the test is triggered by specific change types — caching changes, batch-size changes, dependency upgrades to core runtime libraries, concurrency changes, or significant traffic-volume changes — rather than run purely on a fixed calendar schedule. A time-based backstop (for example, checking at least once per quarter for actively changing services) is a reasonable safety net for changes that don't fall cleanly into a defined trigger category.

Can staging environments be configured to catch this before production? Yes, and it's one of the more effective fixes: configuring staging (or a dedicated performance testing environment) to enforce the same resource limits as production, rather than generous or unset limits, turns staging from a functional-only gate into one that can also surface resource-limit regressions before they reach real users.

Is a memory limit set too low the same problem as one set too high? No, and both are worth watching for separately. A limit set too low relative to genuine, healthy usage causes unnecessary restarts even without any code regression, effectively turning normal variance into recurring incidents — a false-positive version of this problem. A limit set too high, or left effectively unbounded, means a real regression can run for a long time consuming excess memory and crowding out other workloads on shared nodes before anyone notices, since nothing forces a failure. The goal of a resource-limit test plan is finding the accurate number for each service, not simply raising every limit until the restarts stop.

Does this apply the same way to CPU limits as it does to memory limits? Not quite, because the enforcement mechanisms differ. A memory limit breach results in the container being killed, which is abrupt and highly visible once you know to look for OOMKilled. A CPU limit breach results in throttling — the container keeps running but slows down during the throttled window, which tends to show up as latency degradation rather than restarts. Both deserve testing attention, but they require different diagnostic approaches and different test assertions: memory testing looks for approaching a hard ceiling, while CPU testing looks for sustained throttling percentages during realistic load.

Resources and Sources

Recent posts

October 2, 2026
FHIR Interoperability Testing: Certified, Not Connected
October 2, 2026
Definition of Done Erosion: Why Standards Quietly Slip
October 2, 2026
Kubernetes Admission Control Testing: A Field Guide