Latency & Throughput
They are not the same thing
Latency is how long one request takes. Throughput is how many requests complete per second. It is entirely possible to improve one while worsening the other: batching often raises throughput by making work wait longer, while adding replicas raises capacity without making an individual request execute faster when the system was not queueing.
The two are joined by Little's Law, which is as close to a physical law as this field gets:
In a stable system, average concurrency = throughput × average time in the system. At 5,000 completed requests per second and 200ms average latency, about 1,000 requests are present on average. If each request holds a database connection for that whole 200ms, a 100-connection pool caps that path near 500 requests per second no matter how much CPU you add.
L = λ × W L concurrent requests in the system λ arrival rate (throughput) W time in the system (latency) 5,000 req/s × 0.2 s = 1,000 concurrent → a 100-connection pool caps you at 500 req/s
3 components2 connections0:00
Recording…