Sharding
Everything else first
Sharding splits one logical dataset across databases, each owning a slice. It is one way past a single writer's storage or throughput ceiling, alongside decomposing data by function or adopting a distributed store, and it is among the most expensive architectural decisions in this curriculum.
Cache aggressively. Add read replicas. Move cold data out. Buy a bigger machine, twice. Each of those buys you time at a fraction of the complexity, and the team that shards at 20% of a large instance's capacity has taken on a permanent cost to solve a problem they did not yet have.
What sharding costs you is not the splitting. It is everything that used to be free: joins across shards, transactions spanning shards, ORDER BY over the whole dataset, unique constraints, and the ability to change your mind about the shard key.
A shard key is expensive to change. Online resharding is possible, but moving ownership while reads and writes continue is among the hardest operations a team can run.
3 components2 connections0:00
Recording…