Certificate Expiry Outages: The Calendar You Ignore
Share this post

Certificate Expiry Outages Are a Calendar Problem, Not a Technology Problem

A platform engineer at a mid-sized company is staring at a dashboard that makes no sense. Checkout is failing. Not slowly, not intermittently — completely, for every customer, across every region at once. The application logs show nothing unusual: no deploy went out, no database migration ran, no traffic spike, no upstream provider posting an incident. The load balancers report healthy backends. The pods are running. And yet every payment request is dying with a TLS handshake failure, as if the internet itself had quietly stopped trusting the company's own domain.

The eventual answer is almost insultingly mundane. Eighteen months earlier, someone had manually issued a certificate for an internal payment-gateway subdomain during a migration, set a two-year validity period because that was the default the CA offered, and never wired it into the automated renewal system the rest of the fleet used. The person who issued it left the company eleven months ago. Nobody inherited the subdomain. Nobody watched it expire. The certificate simply ran out, on schedule, exactly as certificates do, and the outage that followed had nothing to do with a novel failure mode and everything to do with a date on a document that existed the whole time.

This is the pattern behind a surprising share of the worst production outages in modern computing, and it is a pattern precisely because it is not mysterious. A certificate expiry outage is not caused by an attacker, a bug in application logic, or a scaling limit nobody anticipated. It is caused by a known date, known months in advance, that nobody with the authority and the attention to act on it was watching. The same is true of domain and DNS record expiry: the registrar sends the notices, the resolver eventually stops answering, and the company finds out from a customer, a status-page mention, or a spike in support tickets rather than from its own monitoring.

The uncomfortable part is not that this happens to companies with weak engineering practices. It happens, repeatedly, at organizations with strong engineering practices — organizations that run sophisticated CI/CD pipelines, invest in observability, and would never ship a database migration without a rollback plan. The gap is not technical capability. Automated certificate renewal through the ACME protocol has existed since 2015 and is free. Certificate and domain expiry monitoring is a solved, commodity problem with tooling from every major cloud provider. What is missing at most organizations is not the technology to prevent this outage. It is a named owner, a monitored alert with a lead time long enough to act on, and — critically — proof that the automation everyone assumes is running has actually run.

This article works through several real, independently documented incidents where exactly this happened at organizations with real engineering sophistication, and then builds toward a concrete way to find out whether your own organization is exposed to the same failure, before a customer finds out for you.

The Outage That Took Down a Nation's Mobile Network Over a Certificate Nobody Watched

On December 6, 2018, roughly 32 million mobile customers in the United Kingdom lost data connectivity for most of a day. The outage was not confined to one carrier. O2, along with mobile virtual network operators riding on its infrastructure — giffgaff, Sky Mobile, Lycamobile, and Tesco Mobile — went dark simultaneously. The disruption reached beyond consumer inconvenience: Transport for London's real-time bus arrival system stopped working, and some healthcare providers that relied on mobile connectivity for logistics were affected. The same underlying fault also disrupted SoftBank's network in Japan on the same day, because both carriers were running the same software from the same vendor. (Ericsson apologises for O2 network outage — Computing)

The root cause, confirmed by the vendor, was an expired software certificate embedded in a specific version of the SGSN-MME node software — the component of Ericsson's mobile core network responsible for managing device sessions and mobility management. When the certificate expired, the affected nodes could no longer validate the licenses and software components that depended on it, and they began rejecting or dropping sessions at scale. This was not a licensing certificate on a website. It was buried deep in telecom infrastructure, invisible to any conventional external certificate scanner, sitting inside a vendor's proprietary software stack running on carrier equipment. (Why millions of Brits' mobile phones were knackered on Thursday: An expired Ericsson software certificate — The Register)

What makes this incident worth studying is not the scale of the impact, though that scale is real. It is where the certificate lived. This was not a public HTTPS certificate that any of the standard scanning tools — Qualys SSL Labs, uptime monitors, browser warnings — would ever have flagged. It was an internal software-licensing artifact, several layers removed from anything a security team would normally audit, installed by a vendor, running on infrastructure the carrier operated but did not fully control the internals of. Ericsson's own account acknowledged the fault traced to specific software versions and committed to decommissioning them, but the public postmortem never fully explained why an expiry with months of advance warning was not caught by internal processes at a company with Ericsson's engineering resources. That absence of explanation is itself instructive: even organizations capable of building carrier-grade telecom infrastructure can have blind spots in the mundane, unglamorous discipline of tracking what is about to expire inside their own software supply chain.

The lesson generalizes past telecom. Any organization running third-party software, appliances, or vendor-supplied components has certificates it did not issue, cannot see through its own certificate inventory tooling, and has no contractual visibility into unless it specifically asks the vendor how expiry is handled. "We use ACME everywhere" is a true statement about the certificates a team issues itself. It says nothing about the certificates baked into a load balancer's management plane, a hardware security module's firmware, or — as in this case — a mobile core network vendor's licensing layer.

A Marketing Platform's Entire Product Disappeared Over an Unpaid Domain Bill

Certificate expiry is one half of this failure mode. Domain and DNS expiry is the other, and it is arguably more dangerous because the failure is total rather than partial — a domain that stops resolving does not fail one service, it removes the company from the internet's address book entirely.

In July 2017, marketing automation vendor Marketo's own domain, marketo.com, expired and went unregistered for a period of time before recovery. Customers began reporting problems around the start of the business day: the main website was unreachable, customer login and reporting systems failed, and — because Marketo's own marketing emails embed tracking and redirect links that route through marketo.com — every hyperlink, image, and form embedded in millions of already-sent client marketing emails went dead at the same moment. This was not a contained internal failure. It broke live, already-delivered customer communications retroactively, for every company using the platform. (Marketing giant Marketo forgets to renew domain name. Hilarity ensues — The Register)

The domain was down for roughly two and a half hours before an independent domain specialist and Marketo customer, Travis Prebble, noticed the lapse, registered the domain himself for the cost of a standard registration, and handed control back to the company — an act of goodwell from outside the organization that resolved the outage faster than Marketo's own internal process did. Marketo's CEO later attributed the failure to "process errors with auto-renewals as well as human errors" in a public statement, and the company subsequently reviewed its internal domain-management procedures. (What Happened When Marketo's Domain Name Expired? — ThousandEyes; Marketo CEO blames domain renewal blunder on 'human and process error' — The Drum)

Two details are worth sitting with. First, the financial asymmetry: the domain itself likely cost under fifty dollars a year to keep registered, against an outage that disrupted a publicly traded marketing platform's entire customer base and generated news coverage for days. Second, and more relevant to the argument of this article, is that ICANN's Expired Registration Recovery Policy requires registrars to send renewal reminders at fixed intervals — one month before expiration, one week before expiration, and again within five days after expiration if the domain has not been renewed. (5 Things Every Domain Name Registrant Should Know About ICANN's Expired Registration Recovery Policy — ICANN) Marketo, in other words, almost certainly received at least three separate automated warnings before the domain lapsed. The failure was not a lack of signal. It was that the signal went to an inbox or a person no longer positioned to act on it — the exact ownership-gap pattern this article keeps returning to.

When Microsoft Itself Missed a Certificate Renewal

If the previous two incidents suggest this only happens to companies without world-class engineering organizations, the most recent and most illustrative example says otherwise. In late August 2026, Microsoft 365 experienced a multi-day service disruption tracked internally as incident EX1464935, affecting Exchange Online, Microsoft Teams, SharePoint Online, OneDrive for Business, Microsoft Purview, Microsoft Defender XDR, and the Microsoft 365 admin center. Users experienced authentication failures, email delivery delays, and an inability to sign into Outlook on the web. The disruption began on August 31 and Microsoft did not declare full recovery until several days later, on September 3. (Massive Microsoft 365 outage causes auth issues, service failures — BleepingComputer)

Microsoft's own public statements described the cause in deliberately general terms — "an issue within a core authentication configuration used by multiple Microsoft 365 services" and, separately, "a configuration problem that stopped authentication components from deploying as expected." Independent reporting, however, pointed to something more specific: administrators troubleshooting the outage observed internal error messages referencing an expired certificate thumbprint inside Microsoft's authentication infrastructure, which shifted the operational focus toward certificate renewal during the incident response. (EX1464935: M365 Outage Due to Forgotten Certificate Renewal — Born's Tech and Windows World) It is worth being precise about what is confirmed and what is not: Microsoft's own public language never used the word "certificate," and the more specific certificate-expiry explanation comes from independently observed error output and subsequent reporting rather than an official root-cause statement. Multi-day, multi-service outages of this kind typically involve more than a single expired artifact, and a full technical postmortem, if Microsoft publishes one, may describe a more layered chain of failure than "one certificate expired."

What is not in dispute is the shape of the incident: an internal authentication component that many other services silently depended on stopped working correctly, the disruption cascaded across products that had no obvious architectural reason to fail together, and the fix took days rather than minutes — consistent with a certificate or trust-chain problem buried inside shared internal infrastructure rather than a customer-facing configuration change that could simply be rolled back. Whatever the precise mechanism, the incident is instructive for the same reason the Ericsson case is: an organization with some of the deepest engineering resources in the industry experienced a major, multi-day, multi-product outage whose proximate trigger sits in exactly the category this article is about — an internal authentication or trust artifact with an expiry date that stopped being watched closely enough. If this can happen inside Microsoft's own authentication stack, "we're a sophisticated engineering organization" is not, by itself, a defense.

Germany's Internet Broke Over a DNS Signature, Not a Hack

Certificate expiry and DNS expiry are often discussed as the same problem, but DNS failures deserve their own attention because the mechanism is different and, in some ways, more insidious: DNSSEC, the security extension designed to prevent DNS spoofing, introduces its own signing keys and its own rollover schedule, and getting that rollover wrong can take down resolution for an entire namespace even when every individual domain registration is current and paid for.

On the evening of May 5, 2026, DENIC — the registry operator for Germany's .de top-level domain — began distributing faulty DNSSEC signatures across its zone. For roughly three and a half hours, from just before 10 p.m. until just after 1 a.m., resolvers that validate DNSSEC (a meaningful share of global DNS traffic, though not all of it) rejected .de domains as untrustworthy and refused to resolve them. Major sites including Amazon's German storefront, DHL, Steam, Web.de, and eBay's German operations became unreachable for validating resolvers during the incident, alongside mainstream news outlets. DENIC's own statement attributed the fault to DNSSEC signature distribution but did not, at the time of reporting, disclose the precise underlying cause; independent speculation pointed toward a zone-signing-key rollover gone wrong, though DENIC had not confirmed that specific mechanism. ICANN data cited in coverage of the incident put DNSSEC adoption at roughly 3.6 percent of .de registrations — but with roughly 18 million total .de domains, that fraction still represented on the order of 648,000 affected domains. (It's always DNS: Denic says sorry for crashing Germany's internet — The Register)

This incident is a useful counterweight to the "just set up auto-renewal and you're safe" instinct. DNSSEC was specifically designed to make DNS more trustworthy, and the registry running it is a serious, professional operator, not a small company with an under-resourced ops team. The failure was not "somebody forgot to pay for a domain." It was a scheduled cryptographic rollover — an operation that, done correctly, should be invisible, and done incorrectly, breaks an entire country's access to whichever domains depend on the affected zone. The parallel to certificate rotation is direct: any process that periodically replaces a piece of trust material on a schedule — a signing key, a certificate, a trust anchor — carries the same two failure modes. Either the rotation fails to happen and the old material expires, or the rotation happens but is executed incorrectly and the new material is invalid. Both produce the same symptom to an end user: the thing that worked yesterday does not work today, for reasons that have nothing to do with anyone attacking anything.

Why This Keeps Happening at Organizations That Are Otherwise Good at Engineering

Four incidents, four different organizations, four different technical mechanisms — a vendor's embedded software license certificate, an unpaid domain registration, an internal authentication certificate inside a hyperscale cloud platform, and a DNSSEC key rollover at a national registry. None of these organizations were incompetent. What they share is not a skills gap. It is a structural pattern that recurs across companies of very different sizes and sophistication levels.

The Ownership Gap

Certificates and domains are provisioned once, at a point in time, usually by whoever is doing the migration, launch, or integration that requires them. That person configures the artifact correctly, moves on to the next task, and — unless the organization has a deliberate process for it — the artifact has no durable owner. It sits in a DNS zone file, a load balancer configuration, or a vendor portal, functioning correctly and invisibly, for months or years. The person who set it up may leave the company, change teams, or simply forget it exists, because nothing about day-to-day operations reminds them. Ownership, in the sense of "a specific person or team who will notice this before it expires and knows how to renew it," was never formally assigned; it existed only as tribal memory in one person's head, and tribal memory does not survive reorganizations, attrition, or eighteen months of not thinking about the thing.

This is precisely what happened in the Marketo case — a renewal reminder almost certainly arrived, but not to someone positioned to act on it before the deadline — and it is the most plausible explanation for the internal SGSN-MME certificate in the Ericsson case: an artifact embedded deep enough in vendor software that no team's job description included watching it.

Automation Is Assumed Working, Not Verified Working

Modern organizations frequently do have automated renewal in place for their primary public certificates, most often through an ACME client like Certbot, cert-manager on Kubernetes, or a cloud provider's managed certificate service. This is real progress, and it is also where a second, subtler failure hides: automation that has been configured once and never deliberately tested is not proven automation. It is a hopeful assumption running on a schedule.

ACME renewal can fail silently for reasons that have nothing to do with the certificate itself: a DNS-01 challenge fails because a DNS provider's API token used by the automation expired (a credential-expiry problem nested inside a certificate-expiry defense — the two failure modes can compound); an HTTP-01 challenge fails because a load balancer rule changed and the well-known validation path is no longer routed correctly; a rate limit is hit because of an unrelated bug that triggered thousands of renewal attempts; the automation's own service account loses permissions after an unrelated security tightening; or the renewal succeeds but the new certificate is never actually reloaded into the serving process, so the old, expiring certificate stays live in memory indefinitely. In every one of these cases, "we have automated renewal" was true as a statement about intent and false as a statement about outcome, and nobody discovered the gap until the certificate actually expired.

Let's Encrypt's own integration guidance is explicit that renewal failures should be treated as an expected, recoverable event rather than a rare edge case: it recommends implementing "graceful retry logic in your issuing services using an exponential backoff pattern, maxing out at once per day per certificate," specifically because transient failures in the renewal path are common enough to plan for, not anomalies to be surprised by. The same guidance recommends checking the ACME Renewal Information (ARI) endpoint at least twice daily, and — for anyone managing renewal at real scale — issuing in small, staggered batches rather than all at once, "rather than batching up renewals into large chunks," precisely so that a single infrastructure hiccup does not take out an entire fleet's certificates simultaneously. (Integration Guide — Let's Encrypt) This is not exotic advice. It is the developer of the most widely used free certificate authority on the internet telling every implementer, plainly, that renewal automation fails sometimes and that failure needs to be caught, retried, and — above all — surfaced to a human, not silently absorbed.

Monitoring That Waits for the User to Report the Error

The third and most damaging piece of the pattern is reactive monitoring. Many organizations' actual "certificate monitoring" consists of an uptime check that pings a homepage and reports up or down. That check does not know a certificate expires in nine days. It will only fire the moment the certificate has already expired and the TLS handshake starts failing — by which point the outage has already started, customers are already affected, and the team is now debugging under pressure rather than executing a calm, pre-planned renewal a week in advance.

Proactive monitoring is a genuinely different discipline: it means tracking the actual expiry date of every certificate and every domain registration in the estate, alerting at multiple lead times before expiry (a 30-day informational alert, a 14-day escalation, a 7-day page to an on-call engineer), and — this is the part most programs skip — periodically verifying that the alerting pipeline itself is wired correctly, by deliberately testing it against a real or simulated expiry rather than trusting that it would have fired if it had ever needed to. AWS's own guidance on this is direct: AWS Certificate Manager integrates with Amazon CloudWatch specifically to expose a DaysToExpiry metric so that teams can set alarms well ahead of the actual expiration date rather than discovering it through a failed connection, and AWS publishes separate guidance specifically for certificates imported into ACM (which, unlike ACM-issued certificates, are not eligible for automatic renewal and require active monitoring) describing how to build expiration alerts using CloudWatch Events and Lambda. (AWS Certificate Manager now provides certificate expiry monitoring through Amazon CloudWatch — AWS; How to Monitor Expirations of Imported Certificates in AWS Certificate Manager — AWS Security Blog) The distinction AWS draws — issued-and-auto-renewed versus imported-and-manually-tracked — maps directly onto the ownership gap described above: the moment a certificate leaves the fully automated, platform-managed path, whether because it was issued externally, requires special extended-validation handling, or was imported from a legacy system, it silently reverts to needing a human owner and an explicit monitoring rule, and that transition is exactly where organizations lose track of artifacts.

It is worth being explicit about what this article is not covering. Internal service credentials — API keys, database passwords, service-account tokens, and signing keys used for machine-to-machine authentication inside an organization's own systems — are a related but distinct expiry-and-rotation problem, with different ownership patterns and different blast radii, and are addressed separately elsewhere. This article is scoped specifically to the public-facing surface: the TLS certificates and DNS records that determine whether a customer, a partner, or a browser can reach the company at all.

The Technical Mechanics Worth Getting Right

Understanding why this keeps happening only matters if it translates into a correctly designed defense. Two technical foundations are worth being precise about, because getting the mechanics slightly wrong is a common way that "we have automation" quietly becomes false.

How ACME Renewal Is Supposed to Work

The ACME protocol, standardized as RFC 8555 and popularized by Let's Encrypt, automates the entire certificate lifecycle: a client proves control of a domain (typically via an HTTP-01 challenge, which serves a token from a well-known URL path, or a DNS-01 challenge, which publishes a token as a TXT record), the certificate authority issues a certificate based on that proof, and the client is responsible for tracking the certificate's expiry and repeating the process before it runs out. Because Let's Encrypt's certificates carry a 90-day validity window, the guidance is to renew at roughly one-third of that lifetime remaining — around day 60, leaving a wide margin for retries before the certificate actually becomes invalid at day 90. The ACME Renewal Information (ARI) extension, more recently introduced, lets the CA proactively tell clients when it wants them to renew — useful in scenarios like mass revocation events — and Let's Encrypt recommends clients check it at least twice a day. (Integration Guide — Let's Encrypt)

This design is deliberately generous with time. A properly configured ACME client has weeks of slack between "renewal was due" and "certificate has actually expired." That slack is precisely what makes silent renewal failure so avoidable and so embarrassing when it happens anyway: the system is not failing because there was no time to react. It is failing because nothing was watching whether the renewal that had thirty days of margin actually happened at all.

How DNS and Domain Expiry Actually Unfold

Domain expiry is not a single cliff edge; it is a sequence of stages, each with its own recovery cost, and understanding the sequence changes how urgently a team should treat a missed renewal notice. Under ICANN's Expired Registration Recovery Policy, registrars are required to send reminder notices one month before expiration and again one week before, plus at least one further notice within five days after expiration if the domain has not been renewed. (5 Things Every Domain Name Registrant Should Know About ICANN's Expired Registration Recovery Policy — ICANN) After expiration, the domain typically enters a registrar-defined renewal window during which the original owner can still renew it directly, often at standard price. If that window passes and the registry deletes the registration, ICANN policy requires a mandatory 30-day Redemption Grace Period during which the domain can still be recovered by the original registrant — but during this period, DNS resolution for the domain must be disabled, meaning the outage has already started even though recovery is technically still possible, and recovery during this window typically carries a significant redemption fee well above normal renewal cost. Only after the redemption period lapses does the domain become available for anyone to register, which is the point at which recovery stops being a matter of paying a fee and becomes a matter of hoping nobody else wants the name first. (Implementation of Expired Registration Recovery Policy — ICANN)

The practical implication is that a missed domain renewal is rarely instantly catastrophic and permanently unrecoverable — but every day of delay after the notices start converts a free, trivial fix into a costly, urgent one, and the DNS resolution outage itself begins well before the point of no return. Marketo's recovery in roughly two and a half hours was fast precisely because an external party intervened immediately during the earliest, cheapest recovery window; an organization without that kind of luck, discovering the lapse a week later, could easily be looking at a redemption fee and a multi-day outage instead of a same-morning fix.

Certificate Transparency Logs: A Detection Layer Most Teams Never Use

There is one more technical mechanism worth understanding, not because it prevents expiry directly, but because it solves a different and closely related blind spot: knowing that a certificate exists at all. Since 2013, every publicly trusted certificate authority has been required to submit newly issued certificates to Certificate Transparency (CT) logs — public, append-only ledgers built on Merkle tree structures that make tampering or silent removal detectable. When a CA submits a certificate, the log returns a Signed Certificate Timestamp, a cryptographic promise to include that certificate within a defined Maximum Merge Delay, and Chrome and other major browsers now require valid SCTs before they will trust a public certificate at all. (How CT Works — Certificate Transparency)

The practical consequence for this article's argument is significant: every public certificate ever issued for a company's domains, by any team, through any CA, at any point in the company's history, is discoverable through CT logs — including the ones nobody remembers requesting. Free tools built on CT log data, such as crt.sh, let a team query every certificate ever issued for a given domain or subdomain in seconds. This turns the fintech scenario described earlier from a six-hour manual investigation into something that could have been caught in an afternoon, simply by periodically querying CT logs against the company's known domains and flagging any certificate that does not also appear in the internal renewal inventory. CT monitoring does not replace expiry monitoring — it answers a different question, "what certificates exist that we might not know about," rather than "when does the certificate we know about expire" — but the two are complementary, and an inventory built only from internal tooling will always miss what CT log monitoring catches for free: certificates issued outside the standard pipeline, by a departed employee, a shadow-IT integration, or a vendor acting on the company's behalf.

Comparing the Available Approaches to Renewal and Monitoring

Most organizations do not consciously choose a strategy for certificate and DNS lifecycle management. They accumulate one, artifact by artifact, as different teams solve the immediate problem in front of them at different times. The table below lays out the realistic options and where each one actually breaks, which is more useful than ranking them abstractly, because the right answer is usually "several of these in combination, deliberately chosen," not a single row.

Approach How it actually works Setup effort Primary failure mode Where it fits
Manual tracking (spreadsheet or ticket) A person records expiry dates and creates reminders manually Low The person leaves, the spreadsheet goes stale, nobody re-derives it after infrastructure changes Should not be the primary control for anything customer-facing, at any organization size
Registrar/CA email reminders only Relying on the built-in notices from the registrar or certificate authority None Reminders route to a shared inbox, a former employee, or a finance alias nobody in engineering reads A legal backstop, never a functioning control on its own
ACME automated renewal (Certbot, cert-manager, ACM) Client automatically requests, validates, and installs new certificates on a schedule Moderate — requires a working challenge path and a reload mechanism Renewal succeeds but is never verified; challenge path breaks silently (DNS API token expired, load balancer rule changed) and nobody notices until expiry The correct default for all public TLS certificates the organization directly controls
Expiry monitoring with alerting (CloudWatch, external scanners, uptime tooling) Dashboards and alerts track days-to-expiry independent of the renewal mechanism Low to moderate Alert thresholds set too close to expiry to allow time to act; alert routes to a channel nobody monitors; the monitoring tool itself has an incomplete inventory A mandatory companion to automation, not a substitute for it — this is the layer that catches automation silently failing
DNS/domain expiry tracking (registrar-agnostic monitoring) A separate system tracks domain registration expiry across all registrars the organization uses, independent of any single registrar's own reminders Low Domains registered outside the tracked account (a marketing team's own registrar account, an acquired company's legacy domain) are invisible to the tracking system Necessary specifically because domains, unlike certificates, are frequently registered outside the primary infrastructure account
Deliberate renewal-failure testing (forced expiry in staging, alert-firing rehearsal) The team periodically proves the alerting and renewal pipeline actually works by testing it against a real or simulated failure Low, but requires discipline to schedule and actually execute Skipped in practice because nothing forces it to happen; treated as optional rather than as part of the control itself The layer most organizations omit entirely, and the one that would have caught most of the incidents described earlier

The pattern across every real incident in this article is the same: automation or reminders existed in some form, and the missing layer was verification — proof that the automation or the reminder actually reached someone able to act, tested deliberately rather than assumed.

Three Scenarios Where the Assumption Breaks

The following three scenarios are hypothetical illustrations, not real companies or real events, constructed to show how this failure mode plays out differently depending on the shape of the business. Each follows the same structure: the situation as the team understood it, the assumption nobody had actually verified, the technical or organizational cause, the consequence, the decision point the team faced, and the better approach.

A Fintech Payment API: The Certificate That Outlived Its Owner

Initial situation. A payments infrastructure company runs a public API used by dozens of downstream merchants to process card transactions. The core platform uses ACM-issued, auto-renewing certificates for its main API domain. Two years earlier, during an integration with a legacy banking partner that required a specific certificate chain the platform's standard ACM issuance could not produce, an engineer manually obtained a certificate from a different CA for a dedicated partner-facing subdomain and installed it directly on a specialized gateway appliance outside the normal deployment pipeline.

Hidden assumption. The security team's dashboard, which reported "100% of production certificates covered by automated renewal," was built by querying the ACM API and the primary Kubernetes cert-manager namespace. It had no visibility into certificates installed directly on the standalone gateway appliance, so that manually issued certificate simply did not appear in the inventory the dashboard reported against. The dashboard was accurate about everything it could see and silently blind to the one artifact that mattered most.

Technical and organizational cause. The engineer who set up the integration moved to a different team eighteen months later. The two-year certificate validity period meant the expiry date was far enough in the future that it was never on anyone's near-term radar, and because the artifact lived outside the standard inventory, no expiry alert existed for it anywhere.

Consequence. On the certificate's expiry date, the partner-facing API endpoint began rejecting every TLS handshake. Because this was the endpoint used specifically for high-value settlement transactions with the legacy banking partner, the failure did not show up in general uptime monitoring for the main API — it showed up as settlement transactions failing for one specific partner, initially misdiagnosed as a partner-side network issue because the platform's own status page and general API health checks reported everything green.

The decision point. After roughly six hours of cross-team investigation, an engineer manually inspecting the gateway appliance's configuration found the expired certificate. The company then faced a choice: quietly reissue a replacement certificate through the same manual process and move on, or use the incident to force a broader question — how many other artifacts exist outside the inventoried, automated path, and why does the organization's "100% coverage" dashboard have a blind spot large enough for this to happen at all.

The better approach. The organization chose the second path, treating the incident as a forcing function to build a certificate inventory that queries actual TLS endpoints across every known domain and subdomain — including a periodic external scan, independent of internal tooling self-reports — rather than trusting any single system's account of "everything we manage." Any certificate discovered by the external scan that did not also appear in the automated-renewal inventory was treated as a defect requiring immediate ownership assignment, not merely a data-quality issue to clean up later.

A Healthcare Patient Portal: The DNS Record Nobody Remembered Adding

Initial situation. A healthcare technology company operates a patient portal used by clinics to let patients view lab results and message providers. The main application domain is well managed, with automated certificate renewal and standard monitoring. A secondary subdomain was created three years earlier for a single-sign-on integration with one specific hospital system's identity provider, configured as a CNAME record pointing to that hospital's federation service.

Hidden assumption. The team that built the SSO integration assumed that because the DNS record pointed to infrastructure the hospital system controlled, expiry and availability of that endpoint were the hospital's problem, not theirs — an assumption that was true for the hospital's own certificate lifecycle but false for the DNS record itself, which the healthcare technology company's own registrar controlled and which was subject to the company's own domain renewal cycle, entirely separate from anything the hospital managed.

Technical and organizational cause. During an unrelated domain portfolio cleanup, an engineer reviewing a list of subdomains with low traffic flagged the SSO subdomain as apparently unused, based on aggregate traffic dashboards that undercounted its actual usage because SSO redirect traffic was misclassified in the analytics pipeline. The subdomain's DNS record was deleted as part of a routine cleanup, with sign-off based on the (incorrect) traffic data rather than a direct check with the clinic still actively using that specific integration.

Consequence. Patients at the affected clinic attempting to log in through their hospital's SSO flow began receiving a generic authentication error. Because the failure occurred entirely within a redirect chain — the patient's browser was sent to a hospital-controlled identity provider and then back to a portal address that no longer resolved — neither the healthcare company's uptime monitoring (which checked the main domain, not the deprecated-looking subdomain) nor the hospital's own systems (which correctly redirected to an address that had simply stopped existing) detected an anomaly. The failure was first reported by clinic staff fielding confused patient phone calls, more than a day after the DNS record was removed.

The decision point. The technology company had to decide whether "low measured traffic" was ever an adequate sole criterion for removing a DNS record tied to an external integration, given that the underlying analytics pipeline was known to misclassify certain redirect-based traffic patterns.

The better approach. The company implemented a rule that any DNS record or certificate tied to an external partner integration requires a direct confirmation from the partner relationship owner before removal, regardless of what internal traffic dashboards suggest — treating partner-facing DNS as contractually significant infrastructure rather than as an ordinary internal cleanup candidate, and maintaining an explicit, human-reviewed registry of which subdomains map to which external partners, separate from and cross-checked against raw traffic volume.

A B2B SaaS Platform: The Renewal That Ran and Still Failed

Initial situation. A mid-sized B2B SaaS company migrated its certificate management to cert-manager on Kubernetes eighteen months ago, replacing a previous manual process, and has treated the migration as complete and low-risk ever since. The team considers certificate expiry a solved problem and has not revisited it since the initial rollout.

Hidden assumption. The team assumed that because cert-manager logs showed successful renewal events firing on schedule, the certificates being served to actual customer traffic were current. Nobody had verified that the renewed certificate secret was actually being picked up by the ingress controller's running configuration rather than sitting, correctly renewed, in the Kubernetes secret store while the ingress controller continued serving a cached, older version of the TLS configuration from before the last reload.

Technical and organizational cause. A change to the ingress controller's reload behavior, made months earlier as part of an unrelated performance optimization, disabled automatic reload-on-secret-change for one specific ingress class used by a subset of customer-specific vanity domains, without anyone connecting that change to certificate renewal at the time it was made. cert-manager continued to successfully renew and store new certificates for those domains every renewal cycle. The ingress controller simply never picked them up.

Consequence. For the specific class of customer vanity domains behind the affected ingress class, the certificate actually being presented to browsers was the one that existed at the time of the last full ingress controller restart, months earlier — steadily approaching its real expiry date even as cert-manager's dashboard correctly reported "renewed successfully" for the underlying secret. When the original, unrenewed certificate finally expired, a specific subset of enterprise customers using vanity domains began seeing browser certificate warnings simultaneously, while the majority of the platform's standard-domain customers were unaffected, initially making the issue look customer-specific rather than systemic.

The decision point. The engineering team had to decide whether "renewal succeeded" as reported by the certificate management tool was, by itself, sufficient evidence that the correct certificate was in production — or whether the only trustworthy signal was checking the actual certificate being served on the wire.

The better approach. The team added a synthetic check, independent of cert-manager's own internal reporting, that periodically performs a real TLS handshake against every production hostname and compares the certificate's actual expiry date, as observed on the wire, against the expected renewal schedule — treating "what a client actually receives" as the only ground truth, rather than trusting any internal system's self-report that a renewal had occurred.

Who Actually Owns This? An Ownership and Lifecycle Responsibility Matrix

Every incident above traces back to the same underlying question: who was supposed to notice, and did that person or team have both the visibility and the authority to act in time? The matrix below is a starting point for making that assignment explicit rather than implicit, adaptable to a given organization's structure.

Lifecycle stage Typical default owner (often wrong) Better-assigned owner Key failure if unassigned
Initial certificate/domain provisioning Whoever is doing the migration or launch at the time The platform or infrastructure team, with a mandatory entry into a central inventory before go-live Artifact exists outside any inventory from day one
Domain registration renewal (billing) Finance or a procurement function, disconnected from engineering A named engineering owner co-responsible with finance, both receiving renewal notices Registrar notices reach a billing inbox nobody in engineering monitors
Public TLS certificate renewal Assumed to be "automated" with no further ownership Platform/SRE team, accountable for both the automation running and its verification Automation silently breaks and nobody is accountable for noticing
Internal/mTLS/private-CA certificates Whichever team happened to set up the internal service Platform team, with a policy requiring all internal certs to be inventoried identically to public ones Internal certificates live entirely outside public-facing monitoring and scanning
Partner-facing DNS records and integration certificates The team that built the original integration, informally The partner or integration owner, formally, with a documented dependency and a change-approval requirement before removal Records get deleted in routine cleanup without partner confirmation
Vendor-embedded certificates (appliances, licensed software) Assumed to be the vendor's problem entirely A named internal owner responsible for asking the vendor directly how expiry is monitored and escalated, tracked as a vendor-risk item Nobody internally even knows the artifact exists until it fails
Verification that monitoring/automation actually works Nobody, in most organizations A recurring, calendared responsibility (ironically, one of the few places recurring internal scheduling is appropriate) assigned to SRE or QA, executed as a deliberate test The safety net is assumed to work and has never been tested against a real failure

The last row deserves particular emphasis, because it is the row most organizations omit entirely and the one that would have prevented or shortened every incident described in this article. An ownership assignment for "watch the certificate" is necessary but not sufficient. The organization also needs an owner for "prove the watching mechanism works," which is a fundamentally different and much rarer responsibility.

Testing the Safety Net: How to Verify Monitoring and Automation Actually Work

This is the section most certificate-and-DNS guidance skips, and it is the part that separates organizations that talk about this problem from organizations that have actually closed it. Renewal automation and expiry alerting are both software systems, and like any other software system that matters to production reliability, they should be tested — deliberately, on a schedule, against conditions that resemble the real failure they exist to catch. Treating them as "set up once, trust forever" is the same mistake as writing a disaster recovery plan and never running a recovery drill.

A practical program for negative-testing the safety net includes several distinct exercises, each catching a different class of silent failure:

Forced renewal testing in a non-production environment. Periodically issue a certificate with an artificially short validity window in staging, using the same ACME client, the same challenge configuration, and the same deployment pipeline as production, and confirm the full cycle — request, validation, issuance, installation, and reload of the serving process — completes correctly and automatically, without manual intervention. This catches the exact failure mode in the B2B SaaS scenario above, where renewal succeeded technically but the new certificate never reached the serving process.

Alert-firing rehearsal. Rather than trusting that a "certificate expires in 14 days" alert would fire correctly if it ever needed to, deliberately create a test certificate or DNS record with a known near-term expiry and confirm the monitoring system detects it, generates an alert, and routes that alert to a channel a real human is actually monitoring — not just a channel that technically exists. This is a direct test of the alerting pipeline's actual behavior, not an audit of its configuration on paper.

Inventory reconciliation against live infrastructure. Regularly run an external scan of every known production hostname and subdomain, independent of any internal tool's self-reported inventory, and compare the two lists. Any certificate or domain observed live on the wire but absent from the internal inventory is a defect to be resolved immediately, not a data-quality note for later. This is what caught the fintech scenario's manually issued gateway certificate.

Ownership contact verification. For every artifact in the inventory, confirm on a recurring basis that the listed owner is still at the company, still on the relevant team, and still able to act on an alert — a check that costs almost nothing to run and catches the single most common root cause across every real incident in this article: a reminder that technically fired, addressed to someone no longer positioned to receive it.

Partial revocation or expiry game days. For organizations with mature reliability practices already running chaos-engineering-style exercises, extend that practice to include a deliberately expired or revoked certificate in a controlled environment, observed end to end: does the on-call engineer get paged, does the runbook exist and work, and does the team have the access and authority to execute an emergency renewal without waiting on an approval chain that was designed for routine work, not incidents.

None of these exercises require exotic tooling. They require treating the renewal and monitoring pipeline as a system worth testing on purpose, on the same footing as a payment flow or a login flow, rather than as background infrastructure assumed to be self-sustaining.

The Cost Calculus: What This Discipline Actually Costs Versus What an Outage Costs

The investment side of this equation is genuinely small. Automated certificate renewal through an existing ACME client or a cloud provider's certificate manager typically requires configuration effort measured in hours, not weeks, for a team that does not already have it. Expiry monitoring built on a cloud provider's native tooling, such as ACM's CloudWatch integration, is largely a matter of enabling an existing feature and setting an alarm threshold. Domain registration renewal, in the overwhelming majority of cases, costs less per year than a single engineer's time for a single afternoon. None of the defenses described in this article require significant new infrastructure spend.

The outage side of the equation is asymmetric in the other direction, and not primarily because of the direct cost of downtime, though that is real. It is the character of the outage that makes it expensive in ways that are hard to fully anticipate in advance. A certificate or DNS expiry outage tends to be total rather than partial for whatever domain or subdomain it affects — there is no graceful degradation, no reduced-functionality mode, just a hard failure for every user of that endpoint simultaneously. It is also, reputationally, a difficult failure to explain, because "we forgot to renew something" reads very differently to a customer, a board, or a journalist than "we experienced a novel and complex distributed-systems failure." The Marketo, O2/Ericsson, and Microsoft incidents all attracted specifically pointed public commentary along those lines, independent of and beyond their measurable downtime — the failure mode itself became part of the story, not just the outage duration.

There is a second, less visible cost worth naming directly: the diagnostic time. Every scenario described in this article, real and hypothetical, took the responding team meaningfully longer to diagnose than a more conventional outage would have, precisely because certificate and DNS failures often present as something else — a payment gateway issue, an authentication problem, a partner integration fault — and the team has to work backward through several layers of misleading symptoms before arriving at "the certificate expired" or "the DNS record is gone." An organization that has already inventoried its certificates and domains, and that already knows which ones exist and who owns them, cuts that diagnostic time dramatically, because "check the certificate expiry on this specific hostname" becomes an early step in the runbook rather than a conclusion reached after hours of elimination.

A simple illustrative calculation makes the asymmetry concrete. The figures below are a hypothetical worked example for a mid-sized SaaS company, not a real benchmark or industry survey result, and are meant only to show the shape of the comparison, not to be quoted as a statistic. Assume a platform team spends roughly forty engineering hours building out full CT-log-based certificate inventory, layered expiry alerting, and one annual rehearsal of the alerting pipeline, plus a few hours per quarter maintaining it — a reasonable, if approximate, order of magnitude for a team that does not already have this in place. At a fully loaded engineering cost of roughly 150 dollars per hour, that first-year investment lands somewhere in the six-thousand-to-eight-thousand-dollar range, most of it a one-time build cost. Compare that to a single afternoon-length total outage of a customer-facing API: even a conservative estimate of the engineering time spent on incident response, the customer-support load from a spike in tickets, and the calendar time spent on a postmortem and remediation plan will typically exceed that entire first-year investment on its own, before accounting for any customer churn, contractual service-credit obligations, or reputational cost that a public postmortem search — of the kind this article's own research turned up for real companies — tends to attach to this specific failure mode long after the outage itself is resolved. The point of the exercise is not the precision of the numbers; it is that the investment side of this equation is bounded and largely one-time, while the outage side is open-ended and recurring for as long as the underlying gap stays open.

Scaling Complexity: Startups, Scale-Ups, and Enterprises

The shape of this risk changes predictably with organizational scale, and the right level of process investment changes with it.

A startup with a handful of domains and a single Kubernetes cluster can often get most of the way to safety with a well-configured ACME client, a single dashboard, and one clearly named owner — the risk at this stage is less about process sophistication and more about that one owner leaving without a documented handoff, since there is rarely enough redundancy for institutional knowledge to survive a single departure.

A scale-up is where the real danger concentrates, because this is the stage at which the number of domains, subdomains, environments, and third-party integrations grows faster than the organization's process for tracking them does. New customer-facing vanity domains get added for enterprise deals. New staging and preview environments spin up subdomains automatically. New partner integrations bring DNS records and sometimes partner-issued certificates that live outside the standard pipeline entirely, exactly as in the fintech and healthcare scenarios above. This is typically the stage at which the gap between "certificates we manage" and "certificates that exist" opens widest, because the organization has grown past what any one person can hold in their head, but has not yet built the systematic inventory and ownership discipline that a mature enterprise eventually requires by necessity.

A large enterprise, almost paradoxically, is often in a better position on this specific risk than a scale-up, not because its engineers are more careful, but because compliance requirements — SOC 2, ISO 27001, PCI DSS, and similar frameworks — frequently force a certificate and domain inventory into existence as an audit artifact, whether or not anyone would have built one voluntarily. The corresponding enterprise risk is different: sprawl across dozens of business units, acquired companies bringing their own legacy domains and certificate practices, and internal PKI systems (as discussed earlier in the internal-TLS context) that sit entirely outside whatever public-facing monitoring the compliance program was actually designed to cover.

The multiplication effect is worth naming plainly, because it is what makes this a scale-dependent risk rather than a constant one. A ten-person startup might have two or three domains and a handful of certificates, small enough that one engineer can reasonably hold the full picture in their head. By the time that same company has grown to a few hundred engineers, it is common to find dozens of customer-facing vanity domains from enterprise sales deals, a similar number of short-lived preview and staging subdomains created automatically by CI pipelines, several partner-specific integration endpoints each with their own DNS and sometimes their own certificate requirements, and at least one legacy domain inherited from an early rebrand or a small acquisition that nobody has fully decommissioned. None of these individually looks dangerous. Collectively, they represent an inventory that has grown by an order of magnitude or more without the organization's ownership and monitoring processes growing to match, which is precisely the gap where the fintech and healthcare scenarios above took root. The right response to this pattern is not to slow down domain and subdomain creation — that would fight against how modern SaaS companies actually operate — but to make inventory registration a mandatory, automated byproduct of provisioning, rather than a manual step someone has to remember to do separately.

Warning Signs Your Organization Is Exposed

Several observable signals correlate strongly with exposure to this failure mode, independent of company size:

  • Nobody can produce, within a few minutes, a complete list of every domain and subdomain the company controls, without first checking with several different teams.
  • Certificate renewal is described as "automated" by more than one team, but no one can point to when the automation was last deliberately tested against a real or simulated failure.
  • Domain registrar accounts are owned or paid for through a process disconnected from engineering, such that engineering would not be the first to know if a renewal failed.
  • Expiry alerts, where they exist, have a lead time short enough that a delayed response leaves no real margin — an alert firing three days before expiry is a page, not a warning.
  • The organization has been through at least one prior near-miss (a certificate caught expiring with days to spare, a domain renewal noticed only because someone happened to check) and treated it as a one-off rather than as a signal about the underlying process.
  • Recent acquisitions, migrations, or vendor onboarding have introduced domains, subdomains, or certificates that were never folded into the standard inventory and monitoring pipeline.
  • Any certificate or DNS record tied to a specific person's personal knowledge, rather than to a documented, team-owned process, particularly if that person has changed roles or teams recently.

A Certificate and DNS Lifecycle Ownership Checklist

The following checklist is designed to be run as a standing diagnostic, not a one-time exercise — ideally revisited whenever the organization's domain or certificate footprint changes materially, such as after an acquisition, a major partner integration, or a significant infrastructure migration.

  1. Build a live inventory, not a remembered one. Generate the list of every domain, subdomain, and certificate from an actual scan of production DNS and TLS endpoints, not from asking teams to self-report what they believe they manage.
  2. Assign a named owner to every entry, not a team name. "Platform team" is not an owner; a specific role or on-call rotation that will still exist after any one individual leaves is an owner.
  3. Confirm registrar and CA account access is not a single point of failure. At least two people should have legitimate, documented access to renew any given domain or certificate authority account, and that access should be reviewed when either person changes roles.
  4. Verify renewal automation end to end, not just at the API level. Confirm that a renewed certificate is actually reloaded into the serving process, not merely stored correctly, using the synthetic on-the-wire check described earlier.
  5. Set alert lead times that leave room to act calmly. A useful minimum is layered alerting at 30, 14, and 7 days before expiry, with escalating urgency and escalating recipient lists, not a single alert close to the deadline.
  6. Route alerts to a monitored, on-call-backed channel, not an inbox. If the alert would not page someone the way a production incident does, it is not a control; it is a hope.
  7. Deliberately test the alerting pipeline at least annually, using a real or simulated near-expiry artifact, and treat a failure to alert correctly as a production incident in its own right.
  8. Treat vendor-embedded and appliance certificates as a distinct inventory category, and explicitly ask every relevant vendor how their component's internal certificates are monitored and renewed.
  9. Require partner or integration-owner sign-off before removing any DNS record or certificate, regardless of what internal traffic data appears to show.
  10. Review the full inventory and ownership assignments after every acquisition, major migration, or reorganization, since these are the events most likely to orphan an artifact from its owner.

Frequently Asked Questions

Is this really different from general credential rotation? Yes, in scope and in blast radius. Internal service credentials — API keys, database passwords, machine-identity tokens — typically fail one internal integration at a time and are visible primarily to engineering. Public TLS certificates and DNS records fail in front of every customer and every partner simultaneously, with no internal buffer between the failure and the outside world, which is why they warrant treatment as a distinct availability risk rather than folding them into a broader credential-hygiene program.

We use a managed certificate service from our cloud provider. Are we automatically safe? Largely, for certificates the provider both issues and renews automatically, such as an ACM-issued certificate attached to a load balancer it also manages. The exposure shifts to anything outside that fully managed path: certificates imported from elsewhere, certificates on infrastructure the provider doesn't directly control, and — separately — the domain registration itself, which most cloud certificate services do not manage or monitor at all.

How far in advance should expiry alerts fire? Layered rather than singular: an informational signal around 30 days out, an escalation around 14 days, and a paging alert around 7 days, giving the team multiple opportunities to notice and act before the deadline becomes urgent, rather than a single alert close enough to expiry that a delayed response leaves no margin.

Does DNSSEC make this problem worse or better? Both, depending on execution. Done correctly, DNSSEC prevents a real and serious class of DNS spoofing attacks. Done with an untested key-rollover process, as the 2026 DENIC incident illustrates, it introduces an additional scheduled cryptographic event that can take down resolution for an entire namespace even when every underlying domain registration is current, which means DNSSEC-enabled organizations need rollover testing added to their monitoring discipline, not removed from it.

Who should own this inside an engineering organization — security, platform, or SRE? The specific team matters less than that exactly one function is explicitly accountable, with named individuals inside it, rather than the responsibility being implicitly shared across security, platform, and whichever team originally set up a given artifact — the diffusion of ownership across multiple teams, each assuming another team is watching, is a more common root cause than any single team's negligence.

What's the single highest-leverage first step for a team starting from nothing? Run an external scan of every hostname the company is known to control and compare it against whatever internal inventory currently exists. The gap between those two lists is almost always where the real risk concentrates, and it is usually discoverable in an afternoon.

Where This Fits Into Release Reliability

Certificate and DNS expiry sits at an uncomfortable intersection: it is infrastructure risk that only ever manifests as a customer-facing production failure, which means it belongs in the same reliability conversation as deployment testing and release verification, not filed separately under "ops housekeeping." QAtronic's software testing services work with engineering teams to verify that the safety nets protecting production — not just the application code — actually function under the conditions they exist to catch, including automated testing of renewal and alerting pipelines against simulated near-expiry conditions rather than trusting configuration alone. For organizations restructuring how infrastructure risk gets monitored and owned more broadly, that conversation often overlaps with broader DevOps consulting services work on observability and operational ownership. The goal is not to add another dashboard nobody watches; it is to confirm, deliberately and on a schedule, that the dashboards already in place would actually catch the failure before a customer does.

The Decision This Article Is Really About

Every incident described here had a moment, months or years before the outage, when a reasonable-sounding decision was made: default validity periods were accepted without question, a renewal reminder was allowed to route to whatever inbox happened to be listed, an integration's certificate was treated as someone else's problem because someone else's infrastructure was on the other end of it. None of those decisions were reckless in isolation. Collectively, and left unexamined for long enough, they produced outages at organizations that, in every other respect, knew how to run reliable systems.

The core distinction this article has tried to draw is between having a defense and having a verified defense. Automated renewal that has never been tested against a real failure is a belief, not a control. An expiry alert that has never actually fired in anger is a hope, not a safeguard. A domain owned by "the team that set it up originally" is not owned by anyone once that team has moved on. The technology to close every gap described in this piece already exists, is largely free, and is documented by the organizations that build it. What separates the companies that avoid this outage from the companies that end up in a postmortem is whether someone with real authority treats "prove this still works" as a recurring responsibility, rather than assuming that because it worked when it was set up, it will keep working indefinitely.

The question worth carrying back to your own team is not "do we have certificate monitoring." Nearly every organization would answer yes to that question, including the ones that later ended up in an incident review. The question is: when was the last time anyone deliberately tried to break it, on purpose, to find out.

Recent posts

October 2, 2026
FHIR Interoperability Testing: Certified, Not Connected
October 2, 2026
Definition of Done Erosion: Why Standards Quietly Slip
October 2, 2026
Kubernetes Admission Control Testing: A Field Guide