Handling the obvious failures is the easy part: a server goes down, you spin up another. The failures that actually take out entire regions are the ones that aren’t obvious. A zone that’s slow but not dead, dropping some traffic but still passing health checks. An AWS Region spreads infrastructure across multiple Availability Zones with independent power, cooling, and networking, so if one zone experiences a power event, a network partition, or …