Um momento
0xB0Lesson 12 of 15

Containers and Kubernetes

Package services as container images and run them on Kubernetes with Deployments, Services, probes, config and autoscaling.

30 min 7-question quiz 2 code exercises
By the end of this lesson you can
  • Package a service as a container image with a Dockerfile
  • Describe Kubernetes Pods, Deployments, Services, ConfigMaps and Secrets
  • Configure probes, resources and autoscaling, and reason about rolling updates

Twelve services in four languages, each needing different libraries: “works on my machine” at scale. Containers package a service with everything it needs into an image that runs the same everywhere. A Dockerfile describes how to build it:

Dockerfile
1FROM python:3.12-slim
2WORKDIR /app
3COPY requirements.txt .
4RUN pip install --no-cache-dir -r requirements.txt   # cached unless requirements change
5COPY . .
6USER 1000                                            # don't run as root
7EXPOSE 8080
8CMD ["python", "-m", "ordering"]

Running dozens of containers by hand doesn’t scale, so most teams use Kubernetes, a container orchestrator. You declare the desired state in YAML; Kubernetes keeps reality matching it - restarting crashed containers, rescheduling them when a machine dies, rolling out new versions.

ObjectWhat it is
Podone or more containers that run together; the smallest unit
Deploymentkeeps N replicas of a pod running and rolls out new versions
Servicea stable name and IP that load-balances across a Deployment’s healthy pods - built-in service discovery (http://ordering)
ConfigMap / Secretconfiguration and credentials injected as environment variables or files
Ingress / Gatewayroutes outside traffic in
HorizontalPodAutoscaleradds or removes replicas based on metrics
ordering.yaml
1apiVersion: apps/v1
2kind: Deployment
3metadata:
4  name: ordering
5spec:
6  replicas: 3
7  selector:
8    matchLabels: {app: ordering}
9  strategy:
10    rollingUpdate: {maxSurge: 1, maxUnavailable: 0}
11  template:
12    metadata:
13      labels: {app: ordering}
14    spec:
15      containers:
16        - name: ordering
17          image: ghcr.io/galactic-noodle/ordering:1.8.2
18          ports: [{containerPort: 8080}]
19          env:
20            - name: DB_PASSWORD
21              valueFrom: {secretKeyRef: {name: ordering-db, key: password}}
22          resources:
23            requests: {cpu: 250m, memory: 256Mi}
24            limits: {memory: 512Mi}
25          readinessProbe: {httpGet: {path: /ready, port: 8080}}
26          livenessProbe: {httpGet: {path: /healthz, port: 8080}}
27---
28apiVersion: v1
29kind: Service
30metadata:
31  name: ordering
32spec:
33  selector: {app: ordering}
34  ports: [{port: 80, targetPort: 8080}]

Things worth noticing:

  • The image tag is a specific version, never latest, so every deploy is reproducible and rollback is just the previous tag.
  • Readiness gates traffic; liveness triggers restarts (the observability lesson’s distinction).
  • Resource requests tell the scheduler how much CPU and memory to reserve; a memory limit stops one leaky container from starving its neighbors (it gets OOM-killed instead).
  • Rolling updates replace pods gradually: maxSurge: 1, maxUnavailable: 0 adds one new pod at a time and removes an old one only after the new one is ready - zero downtime, at the cost of a little extra capacity during the rollout.
autoscaling.py
1import math
2
3def desired_replicas(current, current_cpu, target_cpu, minimum, maximum):
4    ratio = current_cpu / target_cpu
5    if abs(ratio - 1) <= 0.1:              # within 10%: leave it alone
6        return current
7    return max(minimum, min(maximum, math.ceil(current * ratio)))
8
9print(desired_replicas(3, 90, 60, 2, 10))   # busy Friday night
10print(desired_replicas(3, 63, 60, 2, 10))   # close enough
11print(desired_replicas(6, 10, 60, 2, 10))   # quiet Tuesday morning
Output
5
3
2

Key takeaways

  • Containers package a service and its dependencies into an image that runs anywhere.

  • Kubernetes keeps declared state: Deployments for replicas and rollouts, Services for discovery.

  • Pin image versions, set readiness and liveness probes, and set resource requests and memory limits.

  • Rolling updates and autoscaling (desired = ceil(current × metric / target)) keep services up and right-sized.

Lesson quiz

7 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: simulate microservice patterns in Python

Build small Python simulations of the patterns - routers, sagas, outboxes, circuit breakers, traces - and run them against sample inputs. They run locally in your browser; no servers or containers needed.

Exercise 1

Autoscale like the HPA

+25 XP

The first input line is min max target_cpu and the starting replica count, like 2 10 60 3. Each following line is the average CPU % measured at that moment. For each measurement compute the desired replicas the way Kubernetes’ HorizontalPodAutoscaler does - ceil(current × measured / target), leaving the count unchanged when the ratio is within 10% of 1, clamped to min and max - and print cpu 90% -> 5 replicas (scale up), (scale down) or (no change). The new count becomes the current count for the next line.

  • A Friday night
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Simulate a rolling update

+25 XP

The input is replicas maxSurge maxUnavailable. Simulate a rolling update from all-old to all-new pods. Each step: (1) start new pods while the total stays within replicas + maxSurge and new pods don’t exceed replicas; they become ready immediately; (2) then stop old pods while the number of running pods stays at least replicas - maxUnavailable. Print step 1: old 4, new 1 after each step until every pod is new, then done in N steps.

If both maxSurge and maxUnavailable are 0, print invalid: the update could never make progress.

  • Surge one at a time
  • Faster, with some unavailability
  • Stuck
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Gostou da aula? 😆👍
Apoie nosso trabalho com uma doação: