Kubernetes Admission Control Testing: Guardrails You Never Verified
The following scenario is a hypothetical, illustrative composite built to demonstrate a real and well-documented class of failure; it does not describe any actual QAtronic client or a specific reported incident.
A platform team at a mid-sized payments SaaS company had a policy they were proud of: no privileged containers, anywhere, in any namespace, full stop. It had been in place for over a year, built on OPA Gatekeeper, reviewed by security, referenced in the company's SOC 2 documentation as a compensating control. When a new engineer asked during onboarding how the company prevented privileged containers from running in production, the answer from three separate people was some version of "we don't allow it — the policy blocks it."
Eleven months after the policy went live, during a genuinely unrelated incident (a control-plane upgrade that triggered a rolling restart of the Gatekeeper webhook pods at the same moment a downstream service was failing over), an engineer applied a debug pod with privileged: true set, intending it to run for twenty minutes while they inspected a node-level networking issue. The apply succeeded. Nobody's phone rang. No Slack alert fired. The debug pod stayed running, forgotten, for six weeks, until a routine Gatekeeper audit scan (a completely separate mechanism from the admission webhook that was supposed to have blocked it in the first place) flagged the constraint violation in its next periodic pass and someone finally noticed.
Nothing about this was a misconfiguration in the way that phrase usually gets used. The Gatekeeper constraint was written correctly. The Rego logic was sound. The policy, read on its own, did exactly what its author intended. What failed was an assumption nobody had tested: that the webhook enforcing the policy would always be reachable within its timeout window, and that if it weren't, the default behavior, inherited quietly from Gatekeeper's own default failurePolicy: Ignore setting, was one the team had actually chosen rather than one that had simply never come up.
This is the gap this article is about. Leadership believed the policy blocked privileged containers cluster-wide. The system believed something narrower and more conditional: it blocks privileged containers cluster-wide, provided the webhook responds inside its timeout, and does something else entirely if it doesn't. Both of those are accurate descriptions of the same YAML file. Only one of them is what got written down.
Why Admission Policies Get Trusted Like Code Used to Be Trusted Before Code Review Existed
Application code at almost any organization mature enough to write a Kubernetes admission policy in the first place goes through a review-and-test ritual before it merges: at minimum a pull request, usually a required reviewer, and typically a CI pipeline that runs unit tests, and increasingly integration tests, before the merge button is even clickable. None of that ritual is perfect — QAtronic has written elsewhere about exactly what code review does and does not verify — but it exists, it is broadly normalized, and skipping it is treated as an exception that needs a reason.
Admission control policy, the YAML, Rego, or CEL that decides which manifests the Kubernetes API server accepts, routinely skips that ritual entirely, for reasons that make sense individually and compound into a real gap collectively.
First, policy code is usually a small fraction of the size of application code, written by a small number of platform or security engineers rather than the broader engineering organization, so it does not accumulate the same institutional pressure toward tooling investment that a codebase touched by fifty engineers does.
Second, a Gatekeeper constraint or Kyverno policy often looks, on the page, like configuration rather than code (a YAML manifest with a spec block), even when its actual logic (a Rego rule, a CEL expression, a Kyverno foreach with nested conditions) is genuinely Turing-complete-adjacent and just as capable of a silent logic bug as a function written in Go or Python.
Third, and most importantly, the policy's job is to prevent bad things from reaching the cluster, and a policy engineer testing that job tends to test the thing they wrote the policy to catch: apply a manifest with a privileged container, confirm it gets rejected, and stop there, because the policy visibly did its job. What almost nobody tests deliberately is the far larger space of things the policy was never asked about: manifests that resemble the blocked pattern closely enough to expose a match clause typo, manifests in a namespace nobody remembered to include in scope, or manifests submitted during the ten seconds the webhook pod was mid-restart.
The result is a policy repository that gets merged and trusted the way application code was trusted before continuous integration and mandatory review existed for it: by author confidence rather than verified evidence. The rest of this piece is about closing that gap with a testing discipline built specifically for what admission control actually is: a runtime enforcement layer inside the Kubernetes API server, with its own toolchain, its own failure modes, and its own reasons a passing test today says nothing about tomorrow.
A Different Control Point Than Infrastructure-as-Code Testing
It is worth addressing directly, rather than glossing over, why this is not simply a Kubernetes-flavored restatement of testing Terraform before apply. The surface-level similarity is real: both are about validating a change to declarative configuration before it takes effect, and both increasingly get described under the umbrella term "policy as code." The mechanics, timing, and failure modes are different enough that treating them as one problem leads a team to under-invest in one while feeling they have covered both.
Infrastructure-as-code testing, the subject of a separate QAtronic article, is a provisioning-time concern. The question it answers is whether a Terraform or OpenTofu plan, computed against a state file and a cloud provider's API, will do what its author intends when applied — whether a resource replacement is actually a safe in-place update or a silent destroy-and-recreate of something stateful, whether the plan respects budget and compliance guardrails, whether it is reversible. The toolchain is Terraform's native terraform test, OpenTofu's tofu test, Terratest, Conftest running against a plan's JSON output, and cloud-specific policy engines like AWS Config rules. The check happens once, at the moment someone runs plan and apply, typically inside a CI pipeline triggered by a pull request, and the object under test is a cloud resource graph: VPCs, load balancers, IAM roles, managed databases.
Admission control testing is a runtime concern, and the object under test is a Kubernetes manifest, not a cloud resource graph. The control point is the API server itself, and it does not run once per pull request — it runs on every single kubectl apply, every Helm release, every reconciliation loop from every controller and operator in the cluster, continuously, for the lifetime of the cluster. A Terraform plan that passes its checks and gets applied is done; the infrastructure exists and the check has finished its job until the next code change. A Kubernetes admission policy is evaluated fresh, from scratch, against every single object creation and update request the API server ever receives, for as long as the cluster runs — which means its failure modes are less about a single bad decision at deploy time and more about behavior under sustained, repeated, unpredictable load: what happens on request number one is not necessarily what happens on request number four hundred thousand, three months later, when the webhook pod is mid-restart during a completely unrelated node drain.
The toolchains do not overlap either. Nothing in Terraform's test framework, Conftest's Rego-over-JSON plan evaluation, or OpenTofu's native testing touches the Kubernetes admission chain at all. Testing Gatekeeper constraints means using gator test, a purpose-built tool from the Gatekeeper project itself. Testing Kyverno policies means using the kyverno test CLI. Testing native ValidatingAdmissionPolicy objects means using kubectl dry-run against a policy bound with --dry-run=server, or the audit-annotation workflow Kubernetes documents directly. None of this is Terraform-adjacent tooling wearing a Kubernetes hat; it is a genuinely separate discipline that happens to share a marketing phrase with infrastructure-as-code testing.
The distinction from QAtronic's existing Kubernetes Resource Limits Need Their Own Test Plan is different in kind. That article is about correctness at the level of a single workload: does this service's declared memory limit match what it actually needs under sustained real load, given caching behavior, batch-size growth, and dependency upgrades. This article is about correctness at the level of the enforcement mechanism itself, applied across every workload in the cluster: does the policy engine, whichever one is in use, actually catch what it claims to catch, consistently, across every namespace, every resource kind, and every failure condition the API server itself can experience. A team could have perfectly tuned resource limits on every workload and still have an admission control layer that fails open under load; the two problems do not overlap and neither testing discipline substitutes for the other.
And the distinction from a zero-trust or service-mesh testing discipline is a difference of layer, not depth. Admission control testing verifies what happens once, at the moment a manifest is submitted to the API server, before any container exists. Zero-trust and service-mesh testing verifies what happens continuously, on the wire, between workloads that are already running: whether an mTLS handshake is actually enforced, whether a NetworkPolicy actually blocks a connection attempt it claims to block. A manifest that admission control correctly approved can still represent a workload whose network behavior a service mesh fails to constrain; these are sequential, not overlapping, controls, and an organization that tests one thoroughly and assumes it has therefore covered the other has simply moved the same blind spot to a different layer of its own stack.
What Actually Runs Inside the API Server
Understanding what to test requires understanding, mechanically, what admission control is and is not. Kubernetes documents this precisely: every request to create, update, delete, or connect to an object passes through an admission control chain inside the API server, after authentication and authorization but before the object is persisted to etcd. The chain has two phases. Mutating admission runs first and can modify the object — injecting a sidecar container, setting a default resource request, adding a label. Validating admission runs second, after mutation is complete, and can only accept or reject the request; it cannot change what is being submitted.
There are, as of Kubernetes 1.30 and continuing through the version in general production use in 2026, three broadly distinct mechanisms for implementing custom admission logic, and the distinction matters for testing because each has different tooling, different latency characteristics, and different failure behavior.
The first is a webhook-based external admission controller, the mechanism OPA Gatekeeper and Kyverno both use by default. The API server calls out over HTTPS to a service running inside (or, less commonly, outside) the cluster, waits for a response within a configured timeoutSeconds, and applies a failurePolicy of either Fail or Ignore if that call errors, times out, or the endpoint is unreachable. Kubernetes' own Dynamic Admission Control documentation is explicit that this is a network call with all the failure modes network calls have: DNS resolution failures, TLS handshake failures, pod-level unavailability during a rolling restart, and simple slowness under CPU pressure on the node running the webhook pod.
The second is Kubernetes' native ValidatingAdmissionPolicy, which reached general availability in Kubernetes 1.30 in April 2024. It moves policy evaluation in-process, inside the API server itself, using the Common Expression Language (CEL) rather than a network call to an external service. This eliminates network-related failure modes entirely for the policies it covers, at the cost of a less expressive rule language than Rego and a shorter list of validation scenarios it is designed for — it works well for CEL-expressible constraints (label presence, field value checks, simple cross-field comparisons) and is not a drop-in replacement for the full expressiveness Gatekeeper or Kyverno offer for complex, templated, or multi-resource policies.
The third, functionally a hybrid, is Kyverno's own architecture, which historically ran as a webhook like Gatekeeper but has increasingly added native, YAML-based policy expressions that avoid requiring a separate policy language altogether, alongside optional support for compiling policies down to CEL for use with ValidatingAdmissionPolicy directly, a bridging capability Kyverno's own blog on applying Validating Admission Policies via the Kyverno CLI documents.
The table below is the reference point the rest of this article builds on.
| Dimension | OPA Gatekeeper | Kyverno | Native ValidatingAdmissionPolicy |
|---|---|---|---|
| Enforcement mechanism | External validating (and optional mutating) webhook | External validating/mutating webhook, with optional native CEL compilation | In-process, inside the API server, via CEL |
| Policy language | Rego (via ConstraintTemplates) | Native YAML/JSON policy expressions | CEL expressions |
| Network dependency for evaluation | Yes: every request is a network call to the webhook pod | Yes: every request is a network call to the webhook pod | No: evaluated in-process |
| Unit-testing tool | gator test / gator verify |
kyverno test CLI |
kubectl dry-run plus audit annotations; no dedicated unit-test CLI as of late 2025 |
| Existing-resource (drift) scanning | Audit controller, periodic (default 60s interval) | Background scan (default hourly), populates Policy Reports | Not built in; requires external tooling to re-evaluate existing objects |
| Staged rollout mechanism | enforcementAction: dryrun on a constraint |
failureAction: Audit plus background: true |
validationActions: [Audit] or [Warn] before [Deny] |
| Default failure behavior on engine unavailability | failurePolicy: Ignore for constraint webhook (fail open) by default |
Configurable per policy; commonly deployed with Ignore for availability |
Not applicable: no external dependency to fail |
This table is not decoration. Every testing layer that follows depends on knowing which row of it applies to the policy engine a given organization actually runs, because the test for "what happens if the engine is unavailable" is meaningless for ValidatingAdmissionPolicy (there is no external engine to become unavailable) and is the single most consequential test for a webhook-based Gatekeeper or Kyverno deployment.
Layer One: Unit-Testing Policy Rules Before They Reach a Cluster
The most basic and most frequently skipped layer of admission control testing is the one closest to conventional software testing: does a specific policy rule, evaluated against a specific manifest, produce the outcome its author intended? This sounds trivial and is not, for a reason specific to how these policies get written.
Hypothetical example. Consider a platform team that maintains a Gatekeeper constraint requiring every container in every Deployment to specify both CPU and memory limits. The constraint's match block scopes it to kinds: ["apps/v1/Deployment"] in every namespace except a documented exemption list. Six months after the policy ships, an engineer refactors a batch-processing service from a Deployment to a CronJob for operational reasons entirely unrelated to the policy. CronJobs suit the workload's actual schedule better than a long-running Deployment ever did. The match block, scoped only to Deployment, has nothing to say about a CronJob's pod template. The refactored service now runs with no resource limits whatsoever, and the policy that was supposed to catch exactly this condition never evaluates the manifest at all, because from the API server's perspective, a CronJob and a Deployment are different kinds, and the constraint was never told to care about the former.
Nothing about the Rego logic inside the constraint template is wrong. The failure is entirely in scope, and it is silent by construction: there is no error, no alert, no log line indicating the policy declined to look at this manifest, because from the policy's point of view, an object it was never told to match simply does not exist.
This is precisely the class of bug unit testing against curated fixtures exists to catch, and it is why both major webhook-based engines ship a dedicated CLI for it. Gatekeeper's gator CLI provides gator test, which evaluates a set of Kubernetes object manifests against a ConstraintTemplate and Constraint entirely locally, with no cluster required, and gator verify, which runs structured test suites defined in YAML: a Suite containing one or more Tests, each of which compiles a ConstraintTemplate and instantiates a Constraint, and each Test containing one or more Cases that assert an expected violation count (an exact number, "at least one," or "zero") and, optionally, a regular expression the violation message must match. This means a platform team can write, alongside the CronJob-scoped constraint above, a test case that asserts a CronJob manifest with no resource limits produces the exact violation count the team intends — and if that number is zero when the team expected one, the test fails in CI before the policy ever reaches a real cluster, rather than six months later when a real batch job runs unconstrained in production.
Kyverno's equivalent is the kyverno test CLI command, which runs a declared set of test resources against a declared set of policies and compares the result to an expected outcome file, explicitly designed, per Kyverno's own documentation, for testing "Kubernetes manifests in YAML format against established policies before they reach a live cluster, particularly useful in pull request workflows," and equally for confirming that a policy modification "consistently produces expected outcomes" when Kyverno itself is upgraded or an existing policy is refactored.
Native ValidatingAdmissionPolicy has a thinner dedicated testing story as of this writing. Kubernetes' own tutorial on exploring validating and mutating admission policies demonstrates testing a CEL policy with kubectl apply --dry-run=server, which evaluates the request through the full admission chain, including the policy, without persisting the object: a genuine test, but one that requires a real (or realistic staging) API server rather than a fully offline unit test, and organizations relying heavily on ValidatingAdmissionPolicy should expect to build more of their own CI scaffolding around dry-run calls than teams using Gatekeeper or Kyverno need to.
A concrete illustration of what that scaffolding needs to catch: a ValidatingAdmissionPolicy requiring every container to declare a memory limit might express the rule as a CEL expression along the lines of object.spec.containers.all(c, has(c.resources.limits) && has(c.resources.limits.memory)). This reads as correct and, against a normal Deployment manifest with a single container, behaves correctly. The same expression evaluated against a pod template that includes an initContainers block, used, for instance, to run a database schema migration before the main application container starts, never inspects initContainers at all, because the CEL expression explicitly walks object.spec.containers and nothing in that path touches the separate initContainers array Kubernetes defines alongside it. An init container with no memory limit passes this policy cleanly, not because the policy author intended to exempt init containers, but because CEL, like any expression language, only evaluates exactly the path it is given, and a policy author who tested only against simple single-container manifests would have no reason to notice the gap until an init container without a memory limit actually triggered a node-level resource problem. Testing this deliberately means maintaining, in whatever dry-run scaffolding a team builds around ValidatingAdmissionPolicy, fixture manifests that specifically include initContainers, ephemeral containers, and any other pod-template field the organization's own workloads actually use, not only the fields a first draft of the CEL expression happened to reference.
| Capability | gator test (Gatekeeper) |
kyverno test (Kyverno) |
|---|---|---|
| Runs fully offline, no cluster required | Yes | Yes |
| Structured suite/test/case hierarchy | Yes (Suite → Test → Case) |
Yes (test manifest referencing resources, policies, expected results) |
| Assertion granularity | Exact/at-least-one/zero violation counts, message regex matching | Pass/fail/skip/warn result per resource-policy-rule combination |
| Designed for CI pipeline integration | Yes: nonzero exit code on failure | Yes: documented GitHub Actions example in official docs |
| Covers mutating policy logic | Limited (primarily validating-focused) | Yes, including foreach and context variables |
The practical failure QAtronic sees most often is not the absence of any tests, but a fixture library that mirrors a vendor's bundled example policies (the sample "block privileged containers" and "require labels" manifests included in the Gatekeeper policy library or Kyverno's public policy repository) rather than manifests that reflect an organization's own actual patterns. A vendor's example test suite proves the vendor's example policy works against the vendor's example manifest. It says nothing about whether that same policy, applied against the organization's actual Helm charts, its actual multi-container sidecar patterns, its actual CronJob-heavy batch architecture, or its actual use of initContainers for schema migrations, behaves the same way. Building a fixture corpus from real manifests, pulled from the organization's own git history, its actual production namespaces (sanitized of secrets), and its actual failed-deploy postmortems, is the single highest-leverage step in this layer, and it is also the step most teams skip because it takes real effort while copying a vendor's example test suite takes almost none.
Testing a Mutating Policy Is a Different Assertion Than Testing a Validating One
Everything described so far assumes the policy under test is validating: it either accepts or rejects a manifest, and a test case asserts a violation count. Both Gatekeeper (through its mutation feature) and Kyverno support mutating policies as well — rules that rewrite an incoming object rather than simply approving or blocking it, injecting a default resource request, adding a required label, or rewriting an image reference to point at an internal registry mirror. These deserve a separate testing conversation because the correct assertion is not "did this get rejected" but "did the object come out the other side looking exactly like this," and teams that reuse validating-style test patterns for mutating policies routinely test the wrong thing.
A validating policy test that only confirms a manifest was not rejected tells a team nothing about whether a mutating rule that ran earlier in the chain actually applied the mutation correctly. Kyverno's test framework supports this distinction directly, allowing a test case to assert against the patched resource rather than only a pass/fail/skip result, and the same discipline matters for Gatekeeper's mutation webhooks, where gator test suites can validate the state of an object after a mutating Assign or AssignMetadata policy has run against it.
A brief, hypothetical illustration. A platform team writes a Kyverno mutating rule intended to inject a podAntiAffinity rule into every Deployment lacking one, so that replicas of the same service are spread across nodes by default rather than by accident. The rule is unit-tested for whether it runs without erroring against a sample Deployment — and it passes, because the mutation logic executes cleanly and the resulting object is admitted. What the test never checks is the shape of the resulting podAntiAffinity block itself: a copy-paste error in the CEL-style JMESPath expression used to build the anti-affinity selector produces a syntactically valid but semantically empty selector, one that technically satisfies the schema but constrains nothing. Every deployment continues landing on whichever nodes the scheduler picks by default, exactly as before the policy existed, and the only signal that anything is different is a Kyverno audit log entry confirming the mutation "succeeded" — which it did, in the narrow sense that it ran without error, while doing nothing useful. A test suite built around "did the mutation run without error" instead of "does the resulting object contain the anti-affinity rule the team actually intended" would never catch this, because the failure is entirely in the mutation's output, not its execution.
The practical rule this suggests: every mutating-policy test case should assert against the specific fields the mutation is supposed to produce, not merely against the absence of an error, and that assertion should be written by reading the intended output directly from the requirement the policy exists to satisfy, not by copying whatever the mutation happened to produce on its first successful run.
What a Minimal CI Gate Actually Looks Like
None of this matters if the tests exist only on a developer's laptop. The entire premise of treating policy like application code is that a pull request modifying a constraint, a ClusterPolicy, or a ValidatingAdmissionPolicy binding cannot merge until its own tests pass, the same way a pull request touching application code cannot merge with a failing unit test. A minimal version of this, combining both major webhook-based engines in a single workflow, looks like this:
name: policy-tests
on:
pull_request:
paths:
- "policies/**"
jobs:
gatekeeper-policy-tests:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Install gator
run: |
curl -L -o gator.tar.gz \
https://github.com/open-policy-agent/gatekeeper/releases/latest/download/gator-linux-amd64.tar.gz
tar -xzf gator.tar.gz && sudo mv gator /usr/local/bin/
- name: Run gator verify against fixture suites
run: gator verify ./policies/gatekeeper/suites/... --recursive
kyverno-policy-tests:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Install Kyverno CLI
run: |
curl -L -o kyverno-cli.tar.gz \
https://github.com/kyverno/kyverno/releases/latest/download/kyverno-cli_linux_x86_64.tar.gz
tar -xzf kyverno-cli.tar.gz && sudo mv kyverno /usr/local/bin/
- name: Run kyverno test against fixture directory
run: kyverno test ./policies/kyverno/tests/
This is deliberately minimal. A real pipeline would add caching, pinned tool versions rather than latest, and a required-status-check branch rule tying the merge button itself to these jobs passing, exactly the way a required unit-test check already gates application-code merges in most organizations. The point is not the specific YAML; it is that the entire five-layer discipline this article describes depends on this kind of gate existing at all, since a fixture library and a set of test cases that nobody's merge process actually enforces provides exactly the same protection as no tests at all: protection that exists on disk but not in practice, the precise gap this article opened with.
Layer Two: Testing Where a Policy's Reach Ends
Every admission policy has a boundary — a set of namespaces, resource kinds, or labeled objects it applies to, and everything outside that boundary it does not evaluate at all. Gatekeeper constraints scope this through match blocks and excludedNamespaces; Kyverno policies scope it through match/exclude blocks referencing kinds, namespaces, and label selectors; native ValidatingAdmissionPolicy scopes it through matchConstraints on the policy and matchResources on its binding. In every case, the boundary is expressed as data — a selector, a list, a label key — and data can be wrong in two directions that produce very different consequences.
Over-reach happens when a policy's exemption logic is broader than intended, and the practical failure mode is a security or compliance control that silently does not apply where leadership believes it does. Hypothetical example. An e-commerce platform's Gatekeeper deployment exempts namespaces labeled environment: staging from a strict "no :latest image tag" constraint, on the reasoning that staging environments need fast iteration and the risk of an unpinned image tag matters less there. A new team stands up a pre-production environment for load-testing checkout flow changes ahead of a seasonal sales peak and, following an internal naming convention nobody had connected to the exemption, labels the namespace environment: staging-loadtest. Because the exemption's label selector was configured with a prefix match rather than an exact match (a decision made for legitimate convenience reasons when the exemption was first written, since at the time only namespaces literally named staging existed), the load-testing namespace is silently exempted too. The load-testing environment, which handles synthetic but realistic payment-flow traffic and is treated by the security team as functionally production-adjacent, now runs with unpinned image tags nobody intended to permit there. The gap surfaces only when a security audit, run manually rather than as a continuous check, happens to sample that namespace.
Under-reach happens when a policy's match scope is narrower than the resource kinds it needs to cover, and it is the mechanism behind the CronJob example in the previous section, generalized. It is worth testing deliberately and specifically for every resource kind a policy is meant to constrain, because Kubernetes offers multiple legitimate ways to produce a running pod, and a match block scoped to one of them by default does not automatically cover the others.
| Resource-creation path | Commonly covered by a Deployment-scoped policy? |
Why it gets missed |
|---|---|---|
Deployment → ReplicaSet → Pod |
Yes, if explicitly matched | Default and most commonly tested path |
StatefulSet → Pod |
Often no | Frequently overlooked because it is used less often, so it is exercised less often in testing |
DaemonSet → Pod |
Often no | Rarely the workload type policy authors have in mind when writing the rule |
CronJob → Job → Pod |
Often no | Two levels of indirection from the CronJob object itself; a policy matching Job directly still needs a separate rule from one matching CronJob |
Bare Pod (no controller) |
Often no | Used for debugging, one-off scripts, and manual troubleshooting: exactly the scenarios most likely to bypass normal deploy pipelines and most likely to need the constraint |
Pod created by a custom operator/CRD controller |
Frequently missed entirely | Operators often create pods through their own reconciliation loop rather than a standard controller, and policy authors testing against kubectl apply manifests rarely think to test against operator-generated objects |
A practical scope-testing checklist, built specifically for this layer rather than reused from a generic Kubernetes hardening checklist, looks like this:
- For every policy, list every Kubernetes resource kind that can ultimately produce a pod in the cluster (not just the kind the policy author had in mind when writing it), and confirm explicitly, in a test case, whether the policy matches or ignores each one.
- For every namespace exclusion or label-based exemption, write a test that submits a manifest with a label or name deliberately close to, but not identical to, the exemption pattern, and confirm the policy still applies. If the exemption is a prefix or substring match, this test should specifically try to trigger a false exemption the way the
staging-loadtestexample did. - Maintain the exemption list itself as a reviewed, dated artifact with an owner and an expiration or re-justification cadence, not an append-only list that accumulates one-off carve-outs no one revisits.
- Test policies against manifests generated by any Helm chart, Kustomize overlay, or operator the organization actually uses in production, not only against hand-written YAML — templating engines routinely produce object shapes (extra annotations, injected labels, wrapped
ownerReferences) that differ subtly from what a policy author tested against by hand. - Periodically — not only when a policy is first written — re-run the full scope-testing checklist against the current set of resource kinds actually running in the cluster, since new controllers, operators, and CRDs get adopted over time and quietly expand the space a policy needs to cover.
Layer Three: What Happens When the Policy Engine Itself Is Unavailable
This is the layer most admission control deployments skip entirely, and it is the one the opening scenario of this article turns on. Every webhook-based admission controller has, by necessity, a defined behavior for what the API server does when the webhook does not answer in time, cannot be reached at all, or returns an error, and that behavior is controlled by two settings most teams set once, during initial installation, and never revisit: failurePolicy and timeoutSeconds.
Kubernetes' own Dynamic Admission Control documentation states plainly that failurePolicy: Fail rejects the request if the webhook call fails for any reason, while failurePolicy: Ignore allows the request through. Gatekeeper's own documentation on customizing admission behavior is direct about which one it chooses by default: the constraint-enforcing webhook defaults to Ignore, "prioritizing availability over enforcement," while a separate, narrower webhook that governs the namespace-exemption labels themselves defaults to Fail, specifically to prevent an attacker or a mistake from using a namespace label to bypass enforcement even during an outage. This split default is a deliberate, documented trade-off, not an oversight — but it is a trade-off almost no team consciously re-evaluates for their own risk tolerance, and it means that on a webhook outage, by Gatekeeper's own default configuration, every constraint stops being enforced while the narrower anti-bypass protection keeps working.
Gatekeeper's own guidance on failing closed is unusually candid about the risk on the other side of this trade-off: setting failurePolicy: Fail for stronger enforcement introduces the possibility of an "admission deadlock": if every node in a cluster is lost and new nodes cannot be scheduled because the webhook that would need to approve the new node's supporting objects is itself unavailable (because its own pods cannot run without those nodes), the cluster can end up in a state where nothing can be created until an administrator manually removes the webhook configuration, which Kubernetes deliberately does not subject to admission validation itself specifically so this recovery path always remains available.
The practical consequence is that "fail open" and "fail closed" are not simply a security-versus-availability dial an organization sets once. They are two distinct failure modes, each with its own blast radius, and testing for them requires deliberately inducing the failure rather than reasoning about it from documentation.
| Failure mode | Trigger | What actually happens | How to test for it |
|---|---|---|---|
| Fail-open (silent) | Webhook pod unreachable, DNS failure, or response exceeds timeoutSeconds, with failurePolicy: Ignore |
Request is admitted as if the policy did not exist; no error surfaces to the requester | Scale the webhook deployment to zero replicas in a non-production cluster (or a dedicated staging namespace mirroring production policy config), submit a manifest that should be rejected, and confirm both that it is admitted and that this produces a detectable signal (metric, log, alert) rather than pure silence |
| Fail-closed (visible) | Same triggers, with failurePolicy: Fail |
Request is rejected with an admission error, even though the manifest itself may be perfectly compliant | Same fault injection as above, but confirm the requester receives a clear, actionable error rather than an opaque timeout, and that an on-call engineer has a documented, tested emergency bypass procedure |
| Admission deadlock | failurePolicy: Fail combined with total loss of nodes running the webhook, with no surviving control-plane path to recreate them |
Cluster cannot admit new objects of any kind, including the objects needed to bring the webhook back up | Deliberately test node-loss recovery in a non-production cluster with failurePolicy: Fail configured identically to production, and confirm the documented manual bypass (removing or patching the ValidatingWebhookConfiguration) actually works under time pressure, not just on paper |
| Timeout misconfiguration | timeoutSeconds on the webhook exceeds the overall API request timeout |
The overall request fails before the webhook's own failure policy is ever invoked, producing a generic client-side timeout rather than the expected admission decision | Load-test the webhook under realistic concurrent request volume and confirm p99 response time stays comfortably under the configured timeout, not just under quiet conditions |
A second hypothetical example, revisiting the opening scenario in more mechanical detail. A fintech company runs Gatekeeper with the default split failurePolicy configuration described above. During a scheduled Kubernetes control-plane minor-version upgrade, the cloud provider's managed control plane briefly increases API server request latency cluster-wide as it drains connections ahead of a rolling restart — a documented, expected behavior during managed upgrades, not a bug. The Gatekeeper webhook pods, running on worker nodes unaffected by the control-plane change, are healthy and responsive under normal conditions, but the elevated end-to-end request latency during the upgrade window pushes several admission requests past the configured timeoutSeconds. With failurePolicy: Ignore on the constraint-enforcing webhook, every one of those specific requests is admitted without constraint evaluation for the several minutes the latency spike lasts. One of those requests happens to be the privileged debug pod from this article's opening. Nothing about Gatekeeper, the constraint, or the debug pod's manifest was faulty. The failure was entirely in an unexercised interaction between a routine infrastructure event and a default configuration setting nobody had deliberately load-tested under realistic latency conditions.
The organizational fix is not necessarily to flip every policy to failurePolicy: Fail. That trade decision depends on the specific policy's actual risk, the team's tolerance for the admission-deadlock scenario, and whether a genuinely critical control (privileged containers, host network access) deserves a different setting than a lower-stakes one (label conventions, naming standards). The fix is to make the choice deliberately, per policy, based on tested behavior under realistic fault conditions, rather than inheriting whatever the tool's installer defaulted to and never revisiting it.
Layer Four: Rolling Out a Policy Change Without Causing an Incident
A policy that is logically correct and properly scoped can still cause an outage the moment it moves from "written" to "enforced," simply because the population of manifests it will now evaluate includes things nobody thought to check against it — legitimate, currently-running workloads that happen to violate a brand-new rule. This is precisely the failure mode a staged rollout exists to catch before it reaches production enforcement, and all three engines this article covers support some version of the same underlying idea: evaluate the policy, report what it would do, and let a human or an automated gate review that report before switching from advisory to blocking.
Native ValidatingAdmissionPolicy exposes this through validationActions, which Kubernetes' own admission policy tutorial documents as a three-step gradient: Audit, which allows the request and records a structured audit annotation without any client-visible effect; Warn, which allows the request but returns a warning message the client (typically kubectl) displays; and Deny, which actually blocks the request. A policy can carry more than one action simultaneously during a transition, running Audit and Warn together, for instance, while the team reviews the volume and nature of audit-logged violations before adding Deny.
Gatekeeper's equivalent is the enforcementAction: dryrun setting on an individual constraint, which evaluates every matching request and records violations without blocking anything, visible through the constraint's own .status field and through the audit controller described in the next section.
Kyverno's equivalent combines failureAction: Audit (the current field name; the older spec.validationFailureAction is being phased out, per Kyverno's own validate-rule documentation) with background: true, which together mean a policy in audit mode both evaluates new incoming requests without blocking them and periodically re-scans existing resources already running in the cluster, surfacing both in Kyverno's Policy Reports.
| Rollout stage | ValidatingAdmissionPolicy | Gatekeeper | Kyverno |
|---|---|---|---|
| Silent observation only | validationActions: [Audit] |
enforcementAction: dryrun |
failureAction: Audit, background: true |
| Visible-but-nonblocking warning | validationActions: [Warn] |
Not natively distinct from dryrun; typically communicated via dashboards built on the audit results | Policy Report entries surfaced via dashboard/alerting; no separate client-side warning mechanism |
| Full enforcement | validationActions: [Deny] |
enforcementAction: deny |
failureAction: Enforce |
The concrete adoption sequence QAtronic recommends for any new or materially changed policy, regardless of which engine is in use, is as follows.
A staged rollout checklist for admission control policy changes:
- Write the policy and its unit-test fixtures together, from the organization's own manifest corpus (not vendor examples), and require both to be reviewed and merged in the same pull request, so a policy can never exist in the repository without its own test coverage.
- Deploy the policy in its audit-only mode (
dryrun,Audit, or equivalent) against the full production cluster — not a staging replica, since the entire point of this stage is to observe the policy's effect against real, currently-running production manifests, which a staging environment cannot fully replicate. - Let the policy run in audit mode for long enough to observe a full deployment cycle for every team whose workloads it touches. A single day is rarely enough, because not every team deploys daily, and a policy that looks clean after one day may still break a workload that only redeploys once a month.
A worked, illustrative sizing exercise. These figures are a hypothetical example built to demonstrate the reasoning, not a benchmark from any real organization. Suppose a platform team is deciding how long to leave a new policy in audit mode before enforcing it, and it profiles its own workloads' deployment cadence: roughly 70 percent of its services redeploy at least weekly through normal CI/CD activity, another 20 percent redeploy roughly monthly, and the remaining 10 percent (mostly quarterly batch jobs, disaster-recovery drill scripts, and a handful of long-lived stateful services nobody has touched recently) redeploy on cycles measured in months rather than weeks. An audit window of three days, chosen because it comfortably covers the 70 percent cohort, would still leave the slowest 10 percent of workloads completely unobserved, meaning the policy could move to full enforcement having genuinely validated compatibility with the bulk of the fleet while remaining a live unknown for exactly the workloads least likely to be redeployed again soon enough to surface a problem before enforcement takes effect. The reasoning this example is meant to illustrate is not a specific number of days; it is that an audit window should be sized against the slowest cohort of workloads large enough to matter, not against the median or the loudest, most frequently deployed services, precisely because the slowest cohort is systematically the one least likely to be exercised by coincidence during a short window and the most likely to include the kind of long-lived, rarely-touched service where an old assumption (like a container image that has always run as root) has had the most time to go unnoticed. 4. Review every audit-logged violation individually and classify each one: a legitimate gap the policy correctly caught (fix the workload), a scope error the policy incorrectly caught (fix the policy's match or exemption logic), or an intentional, justified exception (add it to the reviewed exemption list from Layer Two, with an owner and a reason). 5. Only after that review produces zero unexplained violations does the policy move to enforcing mode, and it moves incrementally where the engine supports it — a subset of namespaces first, using the same scoping mechanism tested in Layer Two, rather than a cluster-wide flip. 6. Immediately after enforcement begins, monitor for a defined window (elevated attention for at least one full business cycle, not a fixed number of minutes) for both false positives (legitimate deploys blocked) and, per Layer Three, silent fail-open events during the switch itself, since a webhook restart to pick up the new enforcement setting is exactly the kind of event that can coincide with a timeout. 7. Record the rollout, including the audit-mode findings and the decision to enforce, somewhere durable and searchable, so the next engineer who asks "why does this policy exist and what does it actually catch" has an answer that does not require reverse-engineering Rego or CEL from scratch.
A third, brief hypothetical example illustrates why step three above specifically warns against a short audit window. A healthcare SaaS company introduces a Kyverno policy requiring every pod to run as a non-root user, deployed with failureAction: Audit and reviewed for a single business day before the team, seeing zero violations, moves it to Enforce. Three weeks later, a quarterly batch reconciliation job (deployed once per quarter, unrelated to the daily deploy cadence every other team observed) fails immediately on its next scheduled run, because its container image, built years earlier by a team that has since moved on, runs as root by design and had simply never been redeployed during the one-day audit window. The policy was correct. The rollout window was too short to observe the actual population of workloads it needed to evaluate, and the incident that resulted was entirely preventable by extending the observation period to cover at least one full cycle of every workload's actual deployment cadence, including infrequent batch and cron-driven jobs, precisely the resource-kind blind spot Layer Two describes, reappearing here as a timing blind spot instead.
Layer Five: Detecting Drift Between the Policy Repository and the Live Cluster
The first four layers all assume that what a team can read in its policy source repository is what is actually running and being enforced inside the cluster. That assumption is not automatically true, and it fails for reasons that mirror, closely, the reasons infrastructure drift happens in Terraform-managed cloud resources, except here the drifted object is the control itself, not the thing being controlled.
Admission policies get patched directly in-cluster during incidents with some regularity, for entirely understandable reasons: a constraint is blocking a legitimate emergency hotfix, an on-call engineer with cluster-admin access edits the constraint's parameters directly with kubectl edit to widen an exemption just enough to let the hotfix through, the incident resolves, and the emergency edit is never backported into the GitOps-managed source repository, because the incident retrospective focuses on the application bug that caused the outage rather than the policy workaround used to route around it during the fire. Weeks or months later, the policy repository and the live cluster describe two different rules, and nobody notices, because nothing about the drifted state produces an error — the cluster is simply enforcing something slightly different from what its own source of truth claims it enforces.
This is a genuinely distinct problem from the drift detection an infrastructure-as-code testing discipline already covers for cloud resources, for one specific reason: a drifted cloud resource (an open security group, a manually resized instance) is itself the thing at risk. A drifted admission policy is the mechanism that was supposed to be preventing risk in the first place, silently no longer doing the job its own source code claims it does: a second-order form of drift, one layer removed from the resource-level drift a general infrastructure drift-detection practice already watches for.
Detecting it requires actively comparing what the live cluster enforces against what the source repository declares, rather than assuming a GitOps sync (Argo CD, Flux, or an equivalent) alone guarantees the two match. A GitOps controller reconciles drift it can see and is configured to correct, but an emergency kubectl edit performed with cluster-admin credentials outside the GitOps pipeline entirely is exactly the kind of change many GitOps configurations are deliberately permissive about during an active incident, and a reconciliation pass that runs on its normal schedule may not immediately revert an emergency change the team intends to keep temporarily, which is itself the moment the change is most likely to be forgotten.
Both major webhook-based engines provide the raw material for drift detection natively, though neither does the full comparison automatically. Gatekeeper's audit controller periodically (every 60 seconds by default) re-evaluates existing cluster resources against every constraint currently loaded, and separately, an operator can query the live ConstraintTemplate and Constraint objects directly from the cluster's API and diff their spec against the same objects' definitions in the source repository — a straightforward scripted comparison, since both are just YAML, but one that has to be built and scheduled deliberately, because neither the audit controller nor Gatekeeper itself performs this specific repo-versus-cluster comparison on its own. Kyverno's Policy Reports work similarly: background scans populate PolicyReport and ClusterPolicyReport objects with the live enforcement outcome, and a scheduled job comparing the live ClusterPolicy objects' spec against the git-tracked source is the same kind of comparison a team has to build rather than receive out of the box.
A fourth hypothetical example. A logistics-software company's platform team discovers, during an unrelated compliance audit eight months after the fact, that a Gatekeeper constraint requiring all LoadBalancer-type services to include an internal-only annotation had been manually patched during an incident to add a namespace-level exemption for a specific team's ingress-testing environment. The exemption was reasonable at the time (that team needed a genuinely public load balancer for three days to validate a partner integration), but the manual patch was never reverted, and the policy repository, which an entirely separate engineer had continued modifying for unrelated reasons over the following eight months, still showed the original, unexempted constraint. Every code review of that repository during those eight months correctly reviewed a policy that the live cluster was not actually running. The gap was invisible to every normal engineering process: code review, CI, even the audit controller itself, since the audit controller correctly reports what the live constraint enforces, which was, accurately, the drifted version, because none of those processes ever compared the live object back to its declared source.
The practical mitigation is a scheduled, automated diff: a CI job, run on a fixed cadence independent of any code change, that pulls every live ConstraintTemplate/Constraint or Kyverno ClusterPolicy object from the cluster's API and compares it field-by-field against the same objects as declared in the source repository, alerting on any discrepancy rather than waiting for an unrelated audit to surface it. This is a modest engineering investment (the comparison itself is a straightforward YAML diff, not a novel testing framework), and it is the layer most commonly missing entirely, because the first four layers all have off-the-shelf tooling built specifically for them, while this one does not, and building "something that doesn't exist yet" competes poorly for engineering time against layers with a CLI already sitting in the tool's own documentation.
What This Costs to Build, and Who Should Own It
None of the five layers above require an exotic engineering investment individually. gator test and kyverno test are both free, open-source CLIs that run in a standard CI pipeline; the fault-injection testing in Layer Three is a scripted scale-to-zero operation against a non-production cluster; the drift-detection job in Layer Five is a scheduled diff. The real cost is organizational: deciding who owns each layer, since admission control sits at the intersection of platform engineering, security, and whichever team's workload the policy actually constrains, and a testing discipline with no clear owner tends to degrade into the same trust-by-default pattern this article opened with.
A reasonable division of ownership, adaptable to team size:
| Layer | Typical owner | Why |
|---|---|---|
| Unit testing (Layer One) | Whoever authors the policy — usually platform or security engineering | Same principle as requiring the author of application code to write its tests; the person closest to the intended logic is best positioned to define the fixtures that would catch a regression in it |
| Scope testing (Layer Two) | Platform engineering, in consultation with the teams whose namespaces or resource kinds are affected | Scope decisions have organizational consequences (which teams get an exemption) that a purely technical owner may not have full visibility into |
| Fail-open/fail-closed testing (Layer Three) | SRE or platform reliability engineering | This is fundamentally a reliability question — behavior under fault conditions — closer to chaos engineering practice than to conventional policy authorship |
| Staged rollout (Layer Four) | Whoever owns the change-management process for the policy repository, typically platform engineering | Rollout discipline is a release-management function, not a testing function per se, even though it depends on the outputs of testing |
| Drift detection (Layer Five) | Platform engineering, with alerting visible to security | This is the layer most likely to be skipped without an explicit owner, since it produces no immediate value until the one day it catches something |
For an organization evaluating whether to build this capability internally or bring in outside help, a useful diagnostic is simply asking, of any admission policy currently running in production: "Show me the test that would catch a scope or fail-open regression in this policy, and tell me when it last ran." A team that has genuinely built this discipline answers with a specific CI job, a specific fixture file, and a specific date. A team that has only configured the policy tends to answer by describing the policy's Rego or YAML instead: a description of intent, not evidence of verified behavior, which is precisely the gap this entire article has been describing.
A Concrete Adoption Sequence
Bringing the five layers together into a single ordered program, rather than five disconnected practices, looks like this for a team starting from an admission control setup that currently has no dedicated testing discipline at all:
- Inventory first. List every admission policy currently enforced in the cluster — not what the policy repository claims should be enforced, but what a live query of
ConstraintTemplate/ConstraintorClusterPolicyobjects against the cluster's actual API shows. This first inventory pass is itself the first drift check, and it is common for it to surface at least one discrepancy immediately. - Build the fixture corpus. Before writing a single new unit test, assemble a library of real manifests (sanitized production examples, past incident postmortems, and representative Helm/Kustomize output) that becomes the shared fixture set every policy's tests draw from, rather than starting each policy's tests from a vendor's bundled examples.
- Retrofit unit tests onto existing policies, prioritized by risk (security-critical constraints like privileged-container blocking first, cosmetic label-convention policies last), using
gator testorkyverno testagainst the fixture corpus from step two. - Run the scope-testing checklist from Layer Two against every existing policy, specifically checking every resource-creation path listed in that section's table, and file the gaps it finds as tracked work rather than treating the exercise as complete once run once.
- Fault-inject the webhook in a non-production cluster configured identically to production's
failurePolicyandtimeoutSecondssettings, confirm the actual behavior matches the intended behavior, and document, per policy rather than as a blanket cluster-wide setting, a deliberatefailurePolicydecision. - Adopt the staged-rollout checklist from Layer Four as the mandatory process for every future policy change, with no exceptions for "small" changes, since the CronJob and non-root-user examples in this article were both small, well-intentioned changes.
- Stand up the drift-detection job from Layer Five as a scheduled CI pipeline, alerting to the same channel the team already uses for production incidents, not a separate low-priority queue that gets checked only during quarterly audits.
- Revisit ownership using the table in the previous section, and confirm each layer has a named owner rather than an assumed one, since the gap in the opening scenario existed for eleven months precisely because everyone believed someone else had verified it.
Where Startups, Scale-Ups, and Enterprises Should Draw the Line
A startup running a single small cluster with a handful of Gatekeeper constraints does not need the full five-layer program on day one, and building it prematurely is itself a form of waste. A reasonable minimum for that stage is Layer One (unit tests for whatever few policies exist, using the vendor CLI directly) and a conscious, documented decision on Layer Three's failurePolicy setting for the one or two policies that actually matter for security, not the full fault-injection exercise, but at minimum reading the documentation and choosing deliberately rather than inheriting a default.
A scale-up running policies across multiple teams' namespaces is the stage where Layer Two's scope testing and Layer Four's staged rollout checklist earn their cost, because this is exactly the stage where a policy written by one team's platform engineer starts affecting workloads that engineer has never personally seen, and the CronJob and non-root-user examples in this article are both scale-up-shaped failures — the organization is large enough that no single person tracks every workload, but not yet large enough to have built the process discipline that would catch the gap automatically.
An enterprise or regulated organization — fintech, healthcare, or any company whose admission control policies get cited in a compliance framework or a customer security questionnaire, the way the opening scenario's payments company cited its privileged-container policy in its own SOC 2 documentation — genuinely needs all five layers, including the drift-detection layer most smaller organizations can reasonably defer, because at that scale an unverified compliance claim is itself a distinct business risk, independent of the underlying technical gap: a security questionnaire answer built on a policy the organization has not actually verified is a factual misstatement waiting to be discovered, not merely a technical debt item.
Questions Engineering Leaders Should Take Back to Their Own Teams
Rather than a generic checklist, the following are specific questions whose answers reveal whether a team has built genuine Kubernetes admission control testing or has only configured policies and assumed the configuration is the same as verified protection. A team with real testing in place answers each with a specific artifact: a CI job, a test file, a date, rather than a description of the policy itself.
- For our most security-critical admission policy, when did its unit tests last run, and what fixture manifests do they actually cover?
- Do we know, for every policy we run, whether its
failurePolicyisFailorIgnore, and did we choose that value deliberately or inherit it from an installer default? - Have we ever deliberately made our admission webhook unavailable in a non-production environment and confirmed what actually happens, rather than reasoning about it from documentation?
- Is there a resource kind, such as a
CronJob, aStatefulSet, or an operator-managedPod, that any of our policies were written withDeploymentin mind but never explicitly tested against? - When a policy was last changed to add an emergency exemption during an incident, was that exemption ever reconciled back into our policy source repository, and how would we know if it wasn't?
- If asked in a customer security questionnaire whether a specific control is enforced across our entire cluster, could we produce a test result proving that, rather than a policy document asserting it?
The Distinction That Actually Matters
The central failure this article describes is not that any particular Gatekeeper constraint, Kyverno policy, or ValidatingAdmissionPolicy object was poorly written. In every example here, the policy logic itself was sound. The failure was always in the space the policy logic never got tested against: a resource kind nobody thought to check, an exemption pattern broader than its author realized, a network timeout nobody deliberately induced before it happened by accident, an audit window too short to see the workload that would eventually violate the rule, and a manual emergency patch nobody reconciled back to its source of truth.
The distinction worth carrying back to any engineering organization running admission control today is this: a policy that has never failed to catch anything is not evidence the policy works. It may simply mean nobody has yet submitted the manifest, during the failure condition, in the namespace, at the resource kind, that the policy was never actually built to handle. Kubernetes admission control testing is the discipline of finding that gap deliberately, on a schedule the organization controls, instead of finding it during an incident, an audit, or — in the least fortunate case — never finding it at all.
Where QAtronic Fits
Building the five-layer testing discipline this article describes (unit tests against a real fixture corpus, scope testing across every resource-creation path, deliberate fault injection against webhook failure behavior, a disciplined staged rollout, and a scheduled drift check between policy source and live cluster state) is a meaningful engineering investment on top of already running Gatekeeper, Kyverno, or ValidatingAdmissionPolicy. QAtronic's Kubernetes consulting services work alongside platform and DevOps teams to build exactly this kind of test coverage for an existing admission control setup, from writing the initial gator test or kyverno test fixture library against an organization's actual manifest patterns through setting up the fault-injection and drift-detection checks most teams never get around to building on their own. The goal is not to replace a platform team's ownership of its own policies, but to help verify that what the policy repository claims is enforced is genuinely what the cluster enforces.
Frequently Asked Questions
Is admission control testing the same thing as running a Kubernetes conformance test suite or a CIS Benchmark scan? No. A CIS Benchmark scan (via a tool like kube-bench) checks whether the cluster's own configuration follows a published hardening baseline: things like API server flags, kubelet settings, and RBAC defaults. It does not test whether a custom admission policy an organization wrote itself correctly enforces the rules that policy claims to enforce, nor does it test that policy's behavior under webhook failure conditions. The two practices are complementary; a cluster can pass every CIS Benchmark check and still have admission policies with the exact gaps this article describes.
We use a managed Kubernetes offering that includes admission control as part of the platform. Do we still need to test our own policies? Yes, for the custom policies layered on top of whatever the managed platform provides by default. A managed control plane may harden its own built-in admission controllers, but any Gatekeeper constraint, Kyverno policy, or ValidatingAdmissionPolicy object an organization's own team authors is exactly as untested as it would be on a self-managed cluster, because the managed provider has no visibility into what that custom policy is supposed to catch or how it is scoped.
How often should the fault-injection test in Layer Three actually run? Tying it only to an annual security review is generally too infrequent, since timeoutSeconds values, webhook replica counts, and cluster latency characteristics all change gradually as infrastructure evolves. Re-running it after any change to webhook deployment configuration, node pool sizing, or Kubernetes version upgrade, and at minimum once per year even absent a specific trigger, is a reasonable default for most organizations past the earliest startup stage.
Does moving fully to native ValidatingAdmissionPolicy eliminate the fail-open/fail-closed problem described in Layer Three? It eliminates the specific failure mode of a network-unreachable webhook, since CEL policies evaluate in-process inside the API server with no external service to become unavailable. It does not eliminate the underlying question of what happens when a policy cannot be evaluated for some other reason — a malformed CEL expression, for instance, has its own error-handling behavior worth testing deliberately — and it does not eliminate the need for Layers One, Two, Four, and Five, all of which apply regardless of which enforcement mechanism a policy uses.
What is the actual difference between Gatekeeper's audit controller and the drift-detection practice in Layer Five? Gatekeeper's audit controller re-evaluates existing cluster objects against the constraints currently loaded in the cluster and reports objects that violate those constraints. It does not compare the constraints themselves against what the policy source repository declares. A constraint that was manually patched in-cluster to be more permissive will be audited correctly against its own (drifted) rules and may show zero violations, even though the source repository describes a stricter rule the cluster is no longer actually running. Layer Five's drift check is specifically about catching that second, subtler gap.
Our policies pass every test in the vendor's own example policy library. Is that sufficient? No, and this is one of the central points of this article. A vendor's bundled example test suite validates that the vendor's example policy behaves as the vendor intended against the vendor's example manifests. It says nothing about whether that policy, or an organization's own variant of it, correctly handles the organization's actual Helm charts, its actual CronJob and StatefulSet usage, or its actual namespace and labeling conventions, which is precisely why Layer One recommends building a fixture corpus from the organization's own real manifests rather than relying on bundled examples alone.