Availability Patterns
Two shapes of fail-over
Active-passive. One node serves the primary workload while a standby prepares to take over. The standby may serve reads or receive warm-up traffic, but its failover capacity is otherwise reserved. Recovery includes detection, promotion, routing, and any cache or connection warm-up.
Active-active. All nodes serve traffic, so load is already spread and there is no promotion from fully idle capacity. The survivors can still need cache warm-up or connection changes, and they must be sized for the failure case: two nodes at 70% each leave one at 140% after a loss.
For stateful systems, a single-primary design accepts writes on one node and promotes a replica during failover; asynchronously replicated writes not yet present on the promoted copy are at risk. Multi-primary accepts writes in more than one place, removing a single promotion path while introducing conflict detection and resolution.
ACTIVE-PASSIVE ACTIVE-ACTIVE ┌────┐ ┌────────┐ ┌────┐ ┌────┐ │ on │ │standby │ │ on │ │ on │ └────┘ └────────┘ └────┘ └────┘ 100% reserved 50% 50% detect + promote + route already serving, but each before full recovery must absorb failure load
Split brain is the failure mode to fear: the heartbeat fails but both nodes are alive, so both believe they are primary and both accept writes. Fencing or a quorum is what prevents it — never a heartbeat alone.
6 components7 connections0:00
Recording…