Curriculum
System Design, end to end.
Tracks build on earlier concepts. Lessons combine concise explanations with predictions, checkpoints, and live systems you can operate where interaction makes the trade-off clearer.
Foundations
How requests travel, and the two numbers that describe every system.
The Request Lifecycle
Follow a single request from a browser to a database and back, hop by hop.
Latency & Throughput
Two numbers people constantly confuse, and the curve that connects them.
Performance vs Scalability
A fast system and a scalable one are different achievements, and you can have either without the other.
DNS
How a name becomes an address, and why it is the most under-appreciated availability dependency you have.
Scaling
Adding capacity, spreading load, and absorbing reads before they cost you.
Horizontal Scaling
Adding replicas works right up until you reach the thing you cannot replicate.
Load Balancing
Spreading traffic across replicas, and what happens when one of them dies.
Reverse Proxies
The layer in front of your servers, and the difference between it and a load balancer.
Caching
Put a cache in front of an overloaded database, then take it away and watch what breaks.
Cache Write Strategies
Cache-aside, write-through, write-behind, and refresh-ahead — four answers to when the cache gets updated.
CDNs & Edge Caching
Move content close to users, and stop most requests before they reach you at all.
Data Systems
Where state lives, how it multiplies, and what it costs to split.
Databases Under Load
Connection pools, query cost, and why the database is almost always the thing that breaks.
Indexes & SQL Tuning
Why a query takes 4ms or 4 seconds, and the handful of things that decide which.
Replication & Stale Reads
Read replicas multiply your read capacity and hand you a consistency problem in exchange.
Federation & Denormalization
Splitting a database by function, and duplicating data on purpose — two moves before sharding.
Sharding
Splitting one database into many — the last resort, and the one that changes everything.
Distributed Systems
Decoupling with queues, surviving partial failure, and agreeing on truth.
Queues & Backpressure
Queues absorb bursts. They do not create capacity, and confusing the two is costly.
Retries & Idempotency
Every network call can fail after succeeding. Designing for that is most of distributed systems.
Circuit Breakers & Bulkheads
Failing fast on a dependency that is already down, so its outage does not become yours.
Production Engineering
Observability, service levels, and operating what you built.
Observability
Metrics, logs, and traces — what each one answers, and why averages hide your worst outages.
SLI, SLO & Error Budgets
Turning reliability from an argument into a number that decides what you work on next.
AI Systems
LLM request architecture, retrieval, and the economics of tokens.
LLM Request Lifecycle
Follow a model request from prompt assembly to the first token, then through every costly token after it.
Structured Outputs
Get machine-usable records out of a probabilistic text generator, and decide what happens when the output does not conform.
Model Routing & Fallback
Put a gateway in front of several models, send each request to the cheapest one that can handle it, and keep serving when a provider stops answering.
Agentic Systems
Tool calling, long-running work, and systems that decide what to do next.