Um momento
0xC0Lesson 13 of 16

Monitoring and alerting

Watch the four golden signals, measure latency with percentiles, set SLOs, and page people only for problems that matter - with burn-rate alerts.

28 min 7-question quiz 2 code exercises
By the end of this lesson you can
  • Monitor services with the four golden signals and percentile latencies
  • Define SLIs and SLOs, and use the error budget to balance speed and reliability
  • Write alerts that are actionable, and alert on how fast the error budget burns

Byte Bakery used to learn about outages from angry posts: “my croissant has been circling my house for an hour”. Deploying many times a day only works if you find out quickly when something goes wrong - ideally before customers do.

Google’s SRE book suggests four golden signals for any user-facing service:

SignalQuestionExample
Latencyhow long do requests take? (successes and failures separately)p99 checkout time
Traffichow much demand is there?orders per second
Errorshow many requests fail?share of 5xx responses
Saturationhow full is the service?CPU, memory, queue length, connection pool use

A typical stack: services expose metrics, Prometheus scrapes and stores them as time series, Grafana draws dashboards, and Alertmanager routes alerts to chat or a pager. Hosted options include Datadog, New Relic and Grafana Cloud; OpenTelemetry standardizes how metrics, logs and traces are produced.

Percentiles, not averages

latency.py
1import math
2
3latencies = [80, 85, 90, 90, 95, 100, 110, 120, 130, 2400]   # milliseconds
4
5def percentile(values, p):
6    ordered = sorted(values)
7    return ordered[math.ceil(p / 100 * len(ordered)) - 1]
8
9print(f"mean {sum(latencies) / len(latencies):.0f} ms")
10print(f"p50 {percentile(latencies, 50)} ms, p90 {percentile(latencies, 90)} ms, p99 {percentile(latencies, 99)} ms")
Output
mean 330 ms
p50 95 ms, p90 130 ms, p99 2400 ms

The mean (330 ms) describes nobody: most requests take about 100 ms, and one customer waited 2.4 seconds. Percentiles tell the real story - p50 is the typical experience, p99 the slowest 1%. At thousands of requests a second, “only 1%” is a lot of hungry people.

SLOs and error budgets

  • A service level indicator (SLI) measures something users care about: “the share of checkout requests that succeed in under 500 ms”.
  • A service level objective (SLO) is the target: “99.9% over 30 days”. (An SLA is a contract with penalties - set it looser than your SLO.)
  • The error budget is what’s left: 0.1% of requests may fail. While there’s budget, ship features fast; when it’s spent, slow down and invest in reliability. 100% is the wrong target - it’s impossibly expensive and users can’t tell 99.99% from 100% through their own Wi-Fi.

Alerts people can trust

Every page should be urgent, actionable and real, and come with a runbook link. Alerting on every CPU spike causes alert fatigue: people start ignoring pages, and then miss the real one.

A robust approach is to alert on the burn rate - how fast you’re spending the error budget. A burn rate of 1 uses the budget exactly over the SLO period; 14.4 would use 2% of a 30-day budget in one hour. Checking a long window and a short one means you page for serious problems, and stop paging soon after they’re fixed. (This setup comes from Google’s SRE workbook.)

alerts.yaml
1groups:
2  - name: checkout-slo
3    rules:
4      - alert: CheckoutErrorBudgetBurn
5        # 99.9% SLO: page when 1h and 5m error rates both exceed 14.4 x 0.1%
6        expr: |
7          (sum(rate(http_requests_total{service="checkout",code=~"5.."}[1h]))
8            / sum(rate(http_requests_total{service="checkout"}[1h]))) > (14.4 * 0.001)
9          and
10          (sum(rate(http_requests_total{service="checkout",code=~"5.."}[5m]))
11            / sum(rate(http_requests_total{service="checkout"}[5m]))) > (14.4 * 0.001)
12        labels:
13          severity: page
14        annotations:
15          summary: Checkout is burning its error budget fast
16          runbook: https://runbooks.bytebakery.example/checkout-errors

Try it

Which golden signal?

Sort each Byte Bakery metric into the golden signal it measures.

0 of 6 sortedScore 0/0
  • “p99 time to load the menu”

  • “Orders placed per second”

  • “Share of payment requests returning 5xx”

  • “Drone dispatch queue length”

  • “Database connection pool 95% in use”

  • “Checkouts that returned “success” but charged nothing”

Key takeaways

  • Watch the golden signals: latency, traffic, errors and saturation.

  • Use percentiles for latency; the mean hides the tail.

  • SLIs measure what users feel, SLOs set targets, and the error budget decides when to slow down.

  • Page only for urgent, actionable problems - burn-rate alerts with a runbook.

Lesson quiz

7 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: automate DevOps chores in Python

Write the small Python tools DevOps teams really build - pipeline runners, plan checkers, metric calculators, scanners - and run them against sample inputs. They run locally in your browser; no servers or cloud accounts needed.

Exercise 1

Compute latency percentiles

+25 XP

The first input line is an SLO: percent threshold_ms, like 95 300 (“95% of requests under 300 ms”). The remaining lines hold latencies in milliseconds, separated by spaces.

Use the nearest-rank method: the p-th percentile is the value at position ceil(p / 100 × n) (counting from 1) in sorted order. Print requests: 20, mean: 182.5 ms (one decimal), p50: 90 ms, p95: 1200 ms, p99: 2400 ms, then SLO 95% under 300 ms: met (96.0% under) or violated (...). “Under” means strictly less than the threshold.

  • Menu service
  • Healthy
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Decide burn-rate alerts

+25 XP

The first input line is the SLO percentage, like 99.9. Each following line is a window: name requests errors for the windows 5m, 30m, 1h and 6h. A window’s burn rate is its error rate divided by the allowed error rate (100% − SLO).

Print each window as 1h: error rate 2.000%, burn rate 20.0x, then the decision: PAGE (1h and 5m burn rates above 14.4x) if both are; otherwise TICKET (6h and 30m burn rates above 6x) if both are; otherwise no alert.

  • Checkout on fire
  • Slow leak
  • Already recovered
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Gostou da aula? 😆👍
Apoie nosso trabalho com uma doação: