The DNS outage that taught us to read our own runbooks
Ninety minutes of partial failure, a runbook that referenced a decommissioned bastion, and what we changed the week after.
The trigger was mundane: I approved a CoreDNS config map rollout with a stale upstream still in it. Recovery was slow because our runbook told us to ssh somewhere that no longer exists.
We now test the first three commands of every critical runbook monthly, in a chaos window, with the on-call engineer who did not write it.
The first symptom was not total outage. About a quarter of our internal lookups failed, mostly from workloads that depended on one external resolver path. That partial failure made triage harder because dashboards looked "degraded" instead of "down", and escalation took longer than it should have.
Our original runbook assumed bastion access to inspect node-level DNS settings. That host was retired during a network hardening project, and the replacement flow lived in a different document. During the incident, we lost twelve minutes switching context between docs and chat transcripts to find the new path.
Once we reached the right control point, rollback was straightforward: restore the previous CoreDNS config map revision and restart affected pods in batches. Service recovered quickly, but we still spent another twenty minutes verifying that cached failures had drained out of client libraries.
The action items were small and strict. Every critical runbook now starts with "last validated" and "validated by" fields. If a command fails during the monthly drill, we fail the drill and fix the doc before closing the ticket.
We also added a dry-run section to DNS changes: check syntax, apply to one non-critical namespace, and confirm success and error-path metrics before full rollout. This added six minutes to deploy time and saved us far more than that the next month when another risky change was caught early.