Monitoring and alerting
Watch the four golden signals, measure latency with percentiles, set SLOs, and page people only for problems that matter - with burn-rate alerts.
- Monitor services with the four golden signals and percentile latencies
- Define SLIs and SLOs, and use the error budget to balance speed and reliability
- Write alerts that are actionable, and alert on how fast the error budget burns
Byte Bakery used to learn about outages from angry posts: “my croissant has been circling my house for an hour”. Deploying many times a day only works if you find out quickly when something goes wrong - ideally before customers do.
Google’s SRE book suggests four golden signals for any user-facing service:
| Signal | Question | Example |
|---|---|---|
| Latency | how long do requests take? (successes and failures separately) | p99 checkout time |
| Traffic | how much demand is there? | orders per second |
| Errors | how many requests fail? | share of 5xx responses |
| Saturation | how full is the service? | CPU, memory, queue length, connection pool use |
A typical stack: services expose metrics, Prometheus scrapes and stores them as time series, Grafana draws dashboards, and Alertmanager routes alerts to chat or a pager. Hosted options include Datadog, New Relic and Grafana Cloud; OpenTelemetry standardizes how metrics, logs and traces are produced.
Percentiles, not averages
1import math
2
3latencies = [80, 85, 90, 90, 95, 100, 110, 120, 130, 2400] # milliseconds
4
5def percentile(values, p):
6 ordered = sorted(values)
7 return ordered[math.ceil(p / 100 * len(ordered)) - 1]
8
9print(f"mean {sum(latencies) / len(latencies):.0f} ms")
10print(f"p50 {percentile(latencies, 50)} ms, p90 {percentile(latencies, 90)} ms, p99 {percentile(latencies, 99)} ms")mean 330 ms p50 95 ms, p90 130 ms, p99 2400 ms
The mean (330 ms) describes nobody: most requests take about 100 ms, and one customer waited 2.4 seconds. Percentiles tell the real story - p50 is the typical experience, p99 the slowest 1%. At thousands of requests a second, “only 1%” is a lot of hungry people.
SLOs and error budgets
- A service level indicator (SLI) measures something users care about: “the share of checkout requests that succeed in under 500 ms”.
- A service level objective (SLO) is the target: “99.9% over 30 days”. (An SLA is a contract with penalties - set it looser than your SLO.)
- The error budget is what’s left: 0.1% of requests may fail. While there’s budget, ship features fast; when it’s spent, slow down and invest in reliability. 100% is the wrong target - it’s impossibly expensive and users can’t tell 99.99% from 100% through their own Wi-Fi.
Alerts people can trust
Every page should be urgent, actionable and real, and come with a runbook link. Alerting on every CPU spike causes alert fatigue: people start ignoring pages, and then miss the real one.
A robust approach is to alert on the burn rate - how fast you’re spending the error budget. A burn rate of 1 uses the budget exactly over the SLO period; 14.4 would use 2% of a 30-day budget in one hour. Checking a long window and a short one means you page for serious problems, and stop paging soon after they’re fixed. (This setup comes from Google’s SRE workbook.)
1groups:
2 - name: checkout-slo
3 rules:
4 - alert: CheckoutErrorBudgetBurn
5 # 99.9% SLO: page when 1h and 5m error rates both exceed 14.4 x 0.1%
6 expr: |
7 (sum(rate(http_requests_total{service="checkout",code=~"5.."}[1h]))
8 / sum(rate(http_requests_total{service="checkout"}[1h]))) > (14.4 * 0.001)
9 and
10 (sum(rate(http_requests_total{service="checkout",code=~"5.."}[5m]))
11 / sum(rate(http_requests_total{service="checkout"}[5m]))) > (14.4 * 0.001)
12 labels:
13 severity: page
14 annotations:
15 summary: Checkout is burning its error budget fast
16 runbook: https://runbooks.bytebakery.example/checkout-errorsTry it
Which golden signal?
Sort each Byte Bakery metric into the golden signal it measures.
“p99 time to load the menu”
“Orders placed per second”
“Share of payment requests returning 5xx”
“Drone dispatch queue length”
“Database connection pool 95% in use”
“Checkouts that returned “success” but charged nothing”
Key takeaways
Watch the golden signals: latency, traffic, errors and saturation.
Use percentiles for latency; the mean hides the tail.
SLIs measure what users feel, SLOs set targets, and the error budget decides when to slow down.
Page only for urgent, actionable problems - burn-rate alerts with a runbook.
Lesson quiz
7 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: automate DevOps chores in Python
Write the small Python tools DevOps teams really build - pipeline runners, plan checkers, metric calculators, scanners - and run them against sample inputs. They run locally in your browser; no servers or cloud accounts needed.
Compute latency percentiles
The first input line is an SLO: percent threshold_ms, like 95 300 (“95% of requests under 300 ms”). The remaining lines hold latencies in milliseconds, separated by spaces.
Use the nearest-rank method: the p-th percentile is the value at position ceil(p / 100 × n) (counting from 1) in sorted order. Print requests: 20, mean: 182.5 ms (one decimal), p50: 90 ms, p95: 1200 ms, p99: 2400 ms, then SLO 95% under 300 ms: met (96.0% under) or violated (...). “Under” means strictly less than the threshold.
- Menu service
- Healthy
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Decide burn-rate alerts
The first input line is the SLO percentage, like 99.9. Each following line is a window: name requests errors for the windows 5m, 30m, 1h and 6h. A window’s burn rate is its error rate divided by the allowed error rate (100% − SLO).
Print each window as 1h: error rate 2.000%, burn rate 20.0x, then the decision: PAGE (1h and 5m burn rates above 14.4x) if both are; otherwise TICKET (6h and 30m burn rates above 6x) if both are; otherwise no alert.
- Checkout on fire
- Slow leak
- Already recovered
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…