At a glance
- Relevance
- Practical value
- Warnings
- None
A single command on a non-production router caused Azure’s global WAN outage because untested SOP changes bypassed safety checks. The real failure was systemic governance gaps, not the engineer who followed the flawed procedure.
Oversimplified incident narratives drive harmful “fixes” like blaming engineers and adding training, missing systemic vulnerabilities that can trigger global outages.
Summary
On January 25, 2023, an engineer ran a command on a non-production router, triggering a global WAN outage impacting all Microsoft cloud services. The simplistic “human error” story masks a far more instructive reality: the incident resulted from a chain of governance gaps, flawed mental models, and unvalidated SOP changes over two months.
The command reset the IGP routing table globally, but only because the SOP was created and modified outside normal change review—no peer review, no emulation testing. The engineer trusted the SOP, believing unsafe commands would be blocked by AAA. The team treated non-production routers differently, skipping listening periods and safety checks.
The real failure lies in systemic deficiencies: inconsistent SOP governance for non-production assets, missing command audits during onboarding, and a culture that equates “following the SOP” with safety. Reactive fixes like adding training or punishing the engineer would have ignored these deeper issues, risking repeat incidents.
Engineering leaders should note that the tactical repairs (e.g., making BGP less sensitive to IGP churn) addressed the immediate failure mode, but the lasting fix required embedding this incident into onboarding training to reshape mental models. The talk underscores the need for blameless analysis that surfaces systemic factors, not just quick closure metrics.
Watch the video
This video is blocked due to your privacy settings. To watch this video, please accept YouTube marketing cookies.
Related breakdowns
How a bank's top team got tired of AI coding agents
A seasoned team that practices TDD and mob programming tried AI for everything. The result: fatigue, reduced flow, and slower delivery. A grounded…
Startup engineering failure is often founder-driven, not technical.
Founders' quirks and hiring shortcuts silently kill startup engineering velocity. David Goldenman explains why trust and alignment matter more than process.
Coding agents are making comprehension the new bottleneck for engineering teams.
Coding agents remove the gatekeeping around code production, but create a new bottleneck: comprehension. The speaker argues that the future belongs to…
Get TL;DW
Too Long; Didn't Watch.
A concise breakdowns of the AI and devtools videos that actually matter for engineering leaders.
Free. Weekly. No hype.
Video and thumbnails remain the property of their respective creators. tldw.news provides editorial analysis, commentary, and discovery links to original content.