Um momento
0xA0Lesson 11 of 15

Observability: logs, metrics and traces

See inside a fleet of services with structured logs, correlation IDs, RED metrics, distributed traces, health checks and SLOs.

30 min 7-question quiz 2 code exercises
By the end of this lesson you can
  • Write structured logs that carry a correlation (trace) ID
  • Monitor services with RED metrics and define SLOs with error budgets
  • Read a distributed trace to find where time goes

A customer on Europa says their order took 9 seconds. In the monolith you’d open one log. Now the request crossed the gateway, Ordering, Menu, Payments and Kitchen - five services, maybe twenty instances. Without observability, you’re guessing.

The three pillars:

  • Logs - what happened, as structured JSON (searchable fields, not prose), each line carrying the request’s trace ID so you can find every line about one request across all services.
  • Metrics - numbers over time. For request-driven services, watch RED: Rate (requests/second), Errors (failed/second), Duration (latency percentiles like p95, p99). For resources, USE: utilization, saturation, errors.
  • Traces - the path of one request through every service, as a tree of timed spans. OpenTelemetry is the vendor-neutral standard for producing all three.
a structured log line
{"time": "2026-10-02T18:30:01.204Z", "level": "warn", "service": "payments",
 "trace_id": "4bf92f3577b34da6", "span_id": "00f067aa0ba902b7",
 "message": "card declined, retrying", "order_id": 42, "attempt": 2}

The trace ID travels with every call, usually in the W3C traceparent HTTP header (and in message headers for events), so each service can tag its logs and spans with it.

Reading a trace

trace.py
1spans = [  # (span id, parent, service and operation, start ms, duration ms)
2    ("a", None, "gateway POST /orders", 0, 940),
3    ("b", "a", "ordering place_order", 5, 925),
4    ("c", "b", "menu get_prices", 10, 60),
5    ("d", "b", "payments charge", 75, 780),
6    ("e", "d", "fraud-check score", 80, 720),
7    ("f", "b", "kitchen create_ticket", 860, 60),
8]
9
10def show(parent, depth):
11    for span_id, span_parent, name, start, duration in spans:
12        if span_parent == parent:
13            bar = " " * (start // 40) + "#" * max(1, duration // 40)
14            label = "  " * depth + name
15            print(f"{label:<28} {duration:>4} ms |{bar}")
16            show(span_id, depth + 1)
17
18show(None, 0)
Output
gateway POST /orders          940 ms |#######################
  ordering place_order        925 ms |#######################
    menu get_prices            60 ms |#
    payments charge           780 ms | ###################
      fraud-check score       720 ms |  ##################
    kitchen create_ticket      60 ms |                     #

The waterfall makes it obvious: most of the 940 ms is the fraud check inside Payments. Without tracing, every team would swear their service was fast - and they’d all be right.

Health checks and SLOs

  • Liveness checks answer “is the process stuck? restart it”. Readiness checks answer “can it serve traffic right now?” - a service warming its cache isn’t ready yet. Keep liveness checks simple: if they depend on the database, a database blip restarts every instance.
  • An SLO (service level objective) is a reliability target users care about: “99.9% of orders succeed over 30 days”. The remaining 0.1% is the error budget - spend it on risky releases, and when it’s gone, slow down and fix reliability.
  • Alert on symptoms users feel (error rate, latency, burning the error budget fast), not on every CPU spike.

Key takeaways

  • Use structured logs with a propagated trace ID; OpenTelemetry standardizes logs, metrics and traces.

  • Watch RED metrics - rate, errors, duration percentiles - for every service.

  • Distributed traces show where time goes across services.

  • Separate liveness from readiness, set SLOs with error budgets, and alert on symptoms.

Lesson quiz

7 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: simulate microservice patterns in Python

Build small Python simulations of the patterns - routers, sagas, outboxes, circuit breakers, traces - and run them against sample inputs. They run locally in your browser; no servers or containers needed.

Exercise 1

Read a trace

+25 XP

Each input line is a span: id parent service start_ms duration_ms (parent - for the root). Print the spans as an indented tree (two spaces per level, children in start order) as payments 780 ms (self 60 ms), where self time is the span’s duration minus the time covered by its direct children (assume children don’t overlap). Then print total: 940 ms and slowest self time: fraud-check (720 ms).

  • Slow order
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Check the error budget

+25 XP

The first input line is the SLO as a percentage (like 99.9). Each following line is day requests failed for one day of the month so far. Print each day’s availability to 3 decimals (day 1: 99.950%), then the totals: availability: 99.912%, budget: 120 failures allowed, 106 used (88%), and a status: healthy (under 50% used), caution (50-99%) or FROZEN: no risky deploys (100% or more).

  • Mid-month
  • Bad week
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Gostou da aula? 😆👍
Apoie nosso trabalho com uma doação: