Retries & Idempotency
A timeout tells you nothing
When a request times out, you have learned exactly one thing: you did not get a response. You have learned nothing about whether the work happened.
Three possibilities, indistinguishable from the caller's side: the request never arrived; it arrived and failed; it arrived, succeeded, and the response was lost. The third is the one that makes naive retries dangerous, and it is not rare.
Exactly-once guarantees have a boundary. Kafka, for example, can atomically commit consumed offsets, state, and produced records inside its transaction model. That does not automatically include an email provider, payment API, or any other external side effect. Across that boundary you still need idempotency, reconciliation, or an explicitly accepted duplicate risk.
client ──request──► server
✗ timeout
Did it happen?
├─ never arrived → retry is correct
├─ arrived, failed → retry is correct
└─ arrived, SUCCEEDED, → retry duplicates it
response lost
You cannot tell which. Design so it does not matter.3 components2 connections0:00
Recording…