On 2 September 2026, zones on our Premium DNS product stopped resolving. We have published a full incident report covering the timeline, where our own response was not good enough, and the commitments we have made as a result, including replacing the provider.
This post is not that.
This post is about the mechanism, because the mechanism is genuinely interesting, and because most of the explanations circulating during the incident were wrong in the same instructive way.
Here’s the short version of the cause: the upstream provider migrated their own domain infrastructure and did not carry over the address records for their four nameserver hostnames. Not the zones. Not the servers. The address records for the names of the servers.
That distinction is the whole story. The DNS data for every affected domain was intact throughout. The machines serving it were up and answering. What disappeared was the ability of any resolver on the internet to work out where to send the question. It is a failure one level above the one everybody assumes when a nameserver stops working, and it behaves completely differently.
In this blog, we outline the resolution chain that broke, why four nameservers gave no protection, and the commands that would have told you the truth in under a minute.
The chain that has to hold
To answer a query for example.com, a resolver walks down a chain of referrals. The root tells it where .com is served. The .com nameservers tell it that example.com is delegated to a set of nameserver hostnames, in this case four names under a single provider domain. Then – and this is the step that matters – the resolver needs an IP address for at least one of those hostnames before it can ask anything.
There are two ways it can get one.
The first is glue. When the parent zone can supply address records for the nameservers alongside the referral, the resolver gets the names and the addresses in the same response and proceeds immediately. This is the common case and it is invisible.
The second is a separate lookup. When glue is absent, expired from cache, or the nameserver names sit outside the zone being delegated, the resolver has to suspend the original query, go and resolve the nameserver hostname as a fresh question from the root down, then come back and resume. This is the out-of-bailiwick case, and it is entirely normal. A large share of the internet’s domains are delegated to nameservers under someone else’s domain, which is what happens the moment you use a managed DNS provider.
Both routes terminate at the same place: address records for those four hostnames. When those records were not carried over during the provider’s migration, both routes went dark simultaneously. Glue had nothing to supply, and the independent lookup had nothing to find. The resolver reached the point in the chain where it needed an address, found no way to obtain one, and gave up.
Why four nameservers protected nothing
Every managed DNS product advertises multiple nameservers, and resellers reasonably read that as redundancy. It is, against one specific failure: a server going down. A resolver that cannot reach ns1 will try ns2, and the outage is absorbed without anyone noticing. That mechanism works well and it’s why single-server DNS failures are rarely newsworthy.
Here, that mechanism provided no protection, because the four nameservers were four names sharing one dependency. They lived under the same provider domain, and their addresses came from the same set of records. When that set went missing, all four names became unresolvable at once. There was no healthy member of the set to fail over to, because the failure was not in any member. It was in the thing all four members were made of.
This is the part worth carrying away from the incident even if you never use this product. Nameserver count is not a measure of redundancy. What matters is whether the nameservers fail independently, and names under one parent domain, maintained as one set of records, by one operator, in one migration, do not fail independently. Four of them is one point of failure with four labels on it.
DNSSEC, for the same reason, would not have helped. Signing a zone protects you against a forged answer. It has nothing to say about the absence of any answer at all.
What SERVFAIL was actually telling you
Affected queries returned SERVFAIL, which is one of the least informative things DNS can tell you, and during the incident it sent a lot of people looking in the wrong place.
SERVFAIL does not mean the record is missing. That is NXDOMAIN. SERVFAIL means the resolver was unable to complete resolution and cannot say why in a way that fits in a response code. From the outside, that looked identical to a dozen unrelated problems. The domain was registered and fine. WHOIS was fine. The zone contents in the control panel were fine, and correct. Nothing a reseller could inspect on their own side was broken, which is exactly why several hours of collective debugging went into infrastructure that was never at fault.
Four commands separate this class of failure from the things it resembles. They are worth knowing before you need them.
Start with the trace, which shows you where the chain stops rather than just that it stopped:
dig +trace example.com
Check what the parent actually says about the delegation, so you know whether the NS set itself is intact:
dig NS example.com @a.gtld-servers.net
Then the diagnostic that would have identified this incident in about four seconds. Ask for the address of the nameserver hostname itself:
dig A ns1.your-dns-provider.example
If that fails while your own zone looks healthy, the problem is not your domain and not your records. It is the provider’s ability to be found. Every affected domain shared that one symptom, and no affected domain had anything else in common.
Finally, if you can get an address for the nameserver from any source, a cached value, documentation, a monitoring system, ask it directly and bypass the naming layer entirely:
dig SOA example.com @198.51.100.53
During this incident, that query answered correctly. The servers were serving the right data for the right zones the entire time. Confirming that is what tells you the outage is in discovery rather than in data, and it changes what you do next, because a data loss and a discovery failure call for opposite responses.
The blast radius, ordered by how long it hurts
DNS failures are usually described as though everything breaks and then everything recovers. In practice the damage is staggered, and the parts that recover slowest are the parts nobody watches during the incident.
Web traffic fails first and recovers first
A, AAAA and CNAME lookups fail, browsers show a resolution error, and the moment resolution returns, so does the site. Recovery is not perfectly simultaneous, because resolvers cache failures as well as successes, and implementations differ in how long they hold a failed lookup. Some clients came back minutes after others.
Mail mostly defers rather than dies
a sending server that cannot resolve MX gets a temporary failure and queues the message, then retries, typically over a period measured in days rather than minutes, though every sender has its own policy and a minority give up early. So the visible mail loss from an outage like this is small, and the invisible consequence is a backlog that lands later, along with whatever your customers concluded when their mail went quiet for an evening.
Email authentication fails in a way that outlives the outage
SPF, DKIM and DMARC are DNS lookups performed by the receiving side, not by you. A receiver holding your MX in cache but unable to resolve your TXT records will process mail it cannot authenticate. Handling varies: some receivers defer, some treat it as unauthenticated and apply policy accordingly. Either way, the record of it sits in receiver reputation systems after your DNS is healthy again, which is the sort of damage you find out about a week later from a deliverability report.
Certificate issuance stops, and that one waits to bite
A CA has to check CAA immediately before issuing, and a CAA lookup that cannot be resolved is not treated as permission. Any issuance or renewal attempted inside the window failed. ACME clients retry, so for most domains this resolves itself invisibly. For a certificate that happened to be close to expiry, it does not, and that failure surfaces days later as an expired certificate with no obvious connection to a DNS incident anyone still remembers.
Why the per-domain workaround did not scale, in TTL terms
The workaround available during the incident was to recreate a zone on our own nameservers and repoint the domain’s delegation. It works. Our incident report is straightforward about the fact that we had no way to do it in bulk, and that this was our gap rather than the provider’s.
What is worth adding here is that the manual labour was not the expensive part. The delegation change was.
Rebuilding a zone by hand is tedious and error-prone, and for a portfolio of a few hundred domains it is a serious amount of work. But once you have done it, you change the NS records at the registry, and then you wait, because resolvers that already hold the old NS set keep using it until it expires. Delegation NS records carry long TTLs by design, commonly 48 hours in .com, and varying by TLD. So the domain you fixed at 18:00 is not reliably fixed at 18:05. It is fixed for resolvers with cold caches, and gradually fixed for everyone else.
In this case the provider’s fix landed at around 02:21. A delegation change made during the evening would in many cases still have been propagating when the underlying problem was already resolved, and then you pay the same propagation cost again to move the zones back. That is the real reason emergency re-delegation is a poor tool for a short vendor outage, and the real reason bulk tooling has to exist before you need it rather than being assembled while the incident is live. It is also why we are building the scripted path now, when the clock is not running.
What we take from this structurally
The provider replacement is covered in the incident report and it is the headline change. The engineering lesson underneath it is the one we would rather other people take away too.
A DNS dependency chain is part of your product’s failure surface, not the vendor’s private business. We knew the product depended on four nameservers. What we had not treated as our problem was that those four names depended on one domain, maintained by one operator, whose own infrastructure migration was not something we would ever see coming or be told about in advance. The dependency was documented. Its correlation was not examined.
So the questions we now ask of an infrastructure vendor are different. Not how many nameservers, but whether the nameserver names resolve through independent paths. Not whether they have redundancy, but what single change on their side can take all of it out at once. Not whether they have an incident process, but whether we can operate around them without one.
Checking your own exposure
None of this is specific to one provider, and the check takes a minute.
To see what your domains are actually delegated to, across whatever portfolio you manage:
while read d; do
printf ‘%s ‘ “$d”
dig +short NS “$d” | sort | tr ‘\n’ ‘ ‘
echo
done < domains.txt
Then look at the output with two questions in mind. Do all the nameservers for a given domain sit under one parent domain, and is that acceptable for what the domain does? And can you resolve each of those nameserver hostnames independently right now?
dig +short A ns1.your-dns-provider.example
For domains where the answer to the first question matters, the fix is nameservers whose names do not share a fate. That means more than one provider, with names under more than one parent domain, which is more work to run and is the honest price of the guarantee.
Domains on Openprovider’s own nameservers were unaffected throughout this incident, and if you want to move zones there, support can help. That is not the point of this post, and it would be a poor lesson to draw from it.
The point is that the four names on your delegation are a dependency worth understanding before somebody else’s migration teaches it to you.
The full incident report, with the timeline and what we are changing, is here.





