Um momento
0x30Lesson 4 of 13

Scale out behind a load balancer

Compare vertical and horizontal scaling and spread traffic across servers.

20 min 6-question quiz 1 code exercise
By the end of this lesson you can
  • Contrast vertical and horizontal scaling.
  • Explain why stateless services are easy to scale out.
  • Choose between common load-balancing algorithms.

Vertical scaling means a bigger machine: simple, but it hits a hardware ceiling and leaves a single point of failure. Horizontal scaling means more machines doing the same job, which adds capacity and redundancy - but only if any server can handle any request. That is why well-designed application tiers are stateless: sessions and data live in shared stores such as a database or Redis, not in server memory.

How load balancers spread traffic

  • Round robin: send each request to the next server in turn. Simple and fair when requests cost about the same.
  • Weighted round robin: bigger servers get proportionally more requests.
  • Least connections: send to the server with the fewest active connections - good when request durations vary.
  • Hashing (by client IP, user id, or with consistent hashing): the same key goes to the same server, useful for cache locality.

A layer 4 balancer routes on IP and port. A layer 7 balancer understands HTTP, so it can route by path or header (for example /api to one pool, /images to another) and terminate TLS. Both use health checks to stop sending traffic to failed instances.

design.py
1servers = ["app-1", "app-2", "app-3"]
2healthy = {"app-1": True, "app-2": False, "app-3": True}
3pool = [server for server in servers if healthy[server]]
4for request_id in range(4):
5    print(f"request {request_id} -> {pool[request_id % len(pool)]}")
Output
request 0 -> app-1
request 1 -> app-3
request 2 -> app-1
request 3 -> app-3

The load balancer must not become the new single point of failure. Production setups run redundant balancers (an active-passive pair with a floating IP, or several behind DNS or anycast). Autoscaling adds and removes application servers as load changes, which only works smoothly because those servers are stateless.

Key takeaways

  • Scale out stateless services; keep state in shared stores.

  • Pick round robin for uniform requests, least connections for uneven ones, hashing for locality.

  • Health checks and redundant balancers keep the tier available.

Lesson quiz

6 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: simulate system design building blocks

Use small Python programs to estimate capacity and simulate caches, load balancers, hash rings, and rate limiters. These exercises run locally in your browser.

Exercise 1

Pick a server by least connections

+25 XP

Read a count n, then n lines of name connections. Print the name of the server with the fewest active connections. If there is a tie, pick the one listed first.

  • Clear winner
  • Tie goes to the first
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Gostou da aula? 😆👍
Apoie nosso trabalho com uma doação: