API gateways and service discovery
Give clients one front door with an API gateway or BFF, and let services find each other with discovery, load balancing and service meshes.
- Explain what an API gateway does: routing, authentication, rate limiting, aggregation
- Choose between one gateway and backends-for-frontends
- Describe service discovery, client- vs server-side load balancing, and what a service mesh adds
The Galactic Noodle Express app shouldn’t need to know there are twelve services, at which addresses. An API gateway is the single front door: clients call api.galacticnoodle.space, and the gateway:
- routes each request to the right service (
/orders/...→ Ordering,/menu/...→ Menu), - handles cross-cutting concerns once: authentication, TLS, rate limiting, request logging, CORS,
- can aggregate: one client request, several service calls, one combined response.
A backend for frontend (BFF) is a gateway tailored to one client type - one for the mobile app, one for the web, one for partner APIs - so each gets exactly the data shape it needs without a one-size-fits-all compromise.
1routes:
2 - prefix: /api/orders
3 service: ordering
4 rewrite: /orders
5 - prefix: /api/orders/history
6 service: order-history
7 rewrite: /history
8 - prefix: /api/menu
9 service: menu
10 rewrite: /dishes
11rate_limit:
12 per_client: 10 # tokens (requests) in the bucket
13 refill_per_second: 2Routing uses the longest matching prefix: /api/orders/history/7 goes to order-history, not ordering, even though both prefixes match. Rate limiting commonly uses a token bucket: each client’s bucket holds up to N tokens and refills at a steady rate; each request spends one, and an empty bucket means 429 Too Many Requests. Bursts are allowed up to the bucket size, but the long-run rate is capped.
How services find each other
Service instances come and go - scaled up on Friday night, replaced during deploys - so their addresses keep changing. Service discovery keeps track:
- A registry (Consul, etcd, Eureka, or Kubernetes itself) knows the healthy instances of each service.
- Server-side discovery: callers use a stable name (
http://menu), and a load balancer or DNS picks an instance. Kubernetes Services work this way. - Client-side discovery: the caller asks the registry for instances and balances between them itself.
A service mesh (Istio, Linkerd) moves all this out of application code into a sidecar proxy next to every instance: discovery, load balancing, retries, timeouts, mutual TLS and telemetry - configured centrally, for every language. Powerful, but one more complex system to run; many teams don’t need one at first.
Key takeaways
An API gateway is the single entry point: routing, auth, rate limiting and aggregation - but no business logic.
Backends-for-frontends tailor a gateway to each client type.
Discovery tracks healthy instances; Kubernetes Services give server-side discovery by name.
A service mesh puts discovery, retries, mTLS and telemetry in sidecar proxies - adopt it when you need it.
Lesson quiz
7 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: simulate microservice patterns in Python
Build small Python simulations of the patterns - routers, sagas, outboxes, circuit breakers, traces - and run them against sample inputs. They run locally in your browser; no servers or containers needed.
Route like a gateway
The input has route lines prefix service rewrite, a line ---, then request lines METHOD path. A route matches when the path equals its prefix or continues with /. Choose the longest matching prefix and replace that prefix with the rewrite: print GET /api/orders/42 -> ordering /orders/42, or GET /api/unknown -> 404 when nothing matches.
- Noodle gateway
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
A token bucket rate limiter
The first input line is capacity refill_per_second. Each following line is a request seconds client (times in non-decreasing order, possibly fractional). Each client has its own bucket, starting full. Before each request, refill the client’s bucket for the time since its last request (up to capacity); then spend one token if available. Print 0.0 zara: 200 (4 left) - tokens left rounded down to a whole number - or 0.4 zara: 429. End with each client’s allowed and rejected counts, alphabetically: zara: 5 allowed, 2 rejected.
- A burst
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…