Observability
Three signals, three questions
Metrics are numeric time series: request rate, error rate, latency percentiles, queue depth. Cheap to store, cheap to query over long windows, and they answer is something wrong, and since when. They cannot tell you which request or which user.
Logs are discrete events with detail. Expensive at volume, and they answer what exactly happened to this one thing. Structured logs — key-value rather than prose — are the difference between grepping and querying.
Traces follow one request across service boundaries, recording the span each service contributed. They answer where did the time go, which in a distributed system is a question neither of the other two can address. The request tracing in this product is the same idea.
A workable default for any service is the RED method: Rate, Errors, Duration. Three metrics, and they cover most of what you need to know about whether something is healthy.
METRICS is something wrong? cheap, aggregate LOGS what happened to X? expensive, detailed TRACES where did the time go? per-request, cross-service
4 components3 connections0:00
Recording…