Skip to content

Redundancy & Failure Domains

Running two copies of something only helps if the thing that could take down the first copy can't also take down the second. A failure domain is the boundary of "things that could fail together" — and redundancy only counts if it crosses that boundary.

flowchart LR Junior["Junior: redundancy that doesn't actually help"] --> Middle["Middle: identifying real failure domains - rack, AZ, region"] Middle --> Senior["Senior: correlated failures across supposedly independent domains"] Senior --> Professional["Professional: failure domain design for real infrastructure"]
flowchart LR subgraph FakeRedundancy["Fake redundancy"] R1["Server 1"] --> SameRack["Same rack,\nsame power circuit"] R2["Server 2"] --> SameRack end subgraph RealRedundancy["Real redundancy"] R3["Server 1: Rack A,\nAvailability Zone 1"] R4["Server 2: Rack B,\nAvailability Zone 2"] end

Choose a level

Level Guide You are done when
Junior Redundancy that doesn't actually help You can identify a "redundant" setup that shares a hidden single point of failure.
Middle Rack, AZ, region You can map redundancy decisions onto real cloud infrastructure failure domain boundaries.
Senior Correlated failures You can explain how failures can correlate across supposedly independent domains.
Professional Failure domain design at scale You can design multi-region redundancy accounting for real, documented correlated-failure incidents.

Practice rule

For any "redundant" pair or set of resources, ask: "what specific, nameable thing (a power circuit, a network switch, a physical building, a cloud provider's control plane) would have to fail to take out ALL of these at once?" If you can't name something sufficiently independent, you don't have real redundancy.