Um momento
0x50Lesson 6 of 15

API gateways and service discovery

Give clients one front door with an API gateway or BFF, and let services find each other with discovery, load balancing and service meshes.

26 min 7-question quiz 2 code exercises
By the end of this lesson you can
  • Explain what an API gateway does: routing, authentication, rate limiting, aggregation
  • Choose between one gateway and backends-for-frontends
  • Describe service discovery, client- vs server-side load balancing, and what a service mesh adds

The Galactic Noodle Express app shouldn’t need to know there are twelve services, at which addresses. An API gateway is the single front door: clients call api.galacticnoodle.space, and the gateway:

  • routes each request to the right service (/orders/... → Ordering, /menu/... → Menu),
  • handles cross-cutting concerns once: authentication, TLS, rate limiting, request logging, CORS,
  • can aggregate: one client request, several service calls, one combined response.

A backend for frontend (BFF) is a gateway tailored to one client type - one for the mobile app, one for the web, one for partner APIs - so each gets exactly the data shape it needs without a one-size-fits-all compromise.

gateway routes
1routes:
2  - prefix: /api/orders
3    service: ordering
4    rewrite: /orders
5  - prefix: /api/orders/history
6    service: order-history
7    rewrite: /history
8  - prefix: /api/menu
9    service: menu
10    rewrite: /dishes
11rate_limit:
12  per_client: 10        # tokens (requests) in the bucket
13  refill_per_second: 2

Routing uses the longest matching prefix: /api/orders/history/7 goes to order-history, not ordering, even though both prefixes match. Rate limiting commonly uses a token bucket: each client’s bucket holds up to N tokens and refills at a steady rate; each request spends one, and an empty bucket means 429 Too Many Requests. Bursts are allowed up to the bucket size, but the long-run rate is capped.

How services find each other

Service instances come and go - scaled up on Friday night, replaced during deploys - so their addresses keep changing. Service discovery keeps track:

  • A registry (Consul, etcd, Eureka, or Kubernetes itself) knows the healthy instances of each service.
  • Server-side discovery: callers use a stable name (http://menu), and a load balancer or DNS picks an instance. Kubernetes Services work this way.
  • Client-side discovery: the caller asks the registry for instances and balances between them itself.

A service mesh (Istio, Linkerd) moves all this out of application code into a sidecar proxy next to every instance: discovery, load balancing, retries, timeouts, mutual TLS and telemetry - configured centrally, for every language. Powerful, but one more complex system to run; many teams don’t need one at first.

Key takeaways

  • An API gateway is the single entry point: routing, auth, rate limiting and aggregation - but no business logic.

  • Backends-for-frontends tailor a gateway to each client type.

  • Discovery tracks healthy instances; Kubernetes Services give server-side discovery by name.

  • A service mesh puts discovery, retries, mTLS and telemetry in sidecar proxies - adopt it when you need it.

Lesson quiz

7 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: simulate microservice patterns in Python

Build small Python simulations of the patterns - routers, sagas, outboxes, circuit breakers, traces - and run them against sample inputs. They run locally in your browser; no servers or containers needed.

Exercise 1

Route like a gateway

+25 XP

The input has route lines prefix service rewrite, a line ---, then request lines METHOD path. A route matches when the path equals its prefix or continues with /. Choose the longest matching prefix and replace that prefix with the rewrite: print GET /api/orders/42 -> ordering /orders/42, or GET /api/unknown -> 404 when nothing matches.

  • Noodle gateway
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

A token bucket rate limiter

+25 XP

The first input line is capacity refill_per_second. Each following line is a request seconds client (times in non-decreasing order, possibly fractional). Each client has its own bucket, starting full. Before each request, refill the client’s bucket for the time since its last request (up to capacity); then spend one token if available. Print 0.0 zara: 200 (4 left) - tokens left rounded down to a whole number - or 0.4 zara: 429. End with each client’s allowed and rejected counts, alphabetically: zara: 5 allowed, 2 rejected.

  • A burst
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Gostou da aula? 😆👍
Apoie nosso trabalho com uma doação: