Reading Latency Percentiles Without Fooling Yourself
By Marcus Hale · June 1, 2026 · Operations
Mean latency figures obscure the critical tail behavior that quantile measurements highlight. An API exhibiting an average response time of fifty milliseconds can still subject five percent of requests to two-second delays; those unfortunate visitors encountering cold starts, cache misses, and saturated shards generate support complaints.
Evaluating quantiles requires understanding distributional nuances. Comparing 95th percentile metrics before and after releases is valid only when underlying traffic compositions remain unchanged; a surge in heavy reporting queries skews tail figures without representing true performance degradation. Similarly, evaluating 99th percentile numbers over narrow 60-second intervals captures fleeting noise rather than structural patterns.
Collect lightweight latency histograms at ingestion boundaries and aggregate metrics centrally over time. Histograms preserve raw distributional shape for arbitrary recalculation, whereas pre-aggregated averages irrevocably destroy underlying statistical details.