Um momento

Learn Distributed Systems

Reason about delay, failure, coordination, and scale across machines.

Start learning 9 lessons · about 2 hours · free
About this track

Learn the core ideas behind reliable services that run across multiple machines. Start with partial failure and retries, then explore clocks, replication, CAP, consensus, partitioning, queues, and recovery. The small Python code panels are illustrative simulations, not networked systems; the track is about system behavior and trade-offs.

Before you start
  • No distributed-systems experience is required.
  • Basic programming familiarity helps with the illustrative Python snippets, but lessons and quizzes do not require running code.
  1. Unit 1 · 0/3 lessons

    Failure and time

    Handle uncertainty, retries, and ordering across machines. Badge: Failure analyst

    1. 0x001Think in distributed systemsLearn what changes when a program depends on multiple machines. 14 min 14 min
    2. 0x102Use timeouts and retries safelyChoose retry policies that recover from transient failures without multiplying work. 14 min 14 min
    3. 0x203Reason about time and orderingDistinguish wall-clock timestamps from causal order. 14 min 14 min
  2. Unit 2 · 0/3 lessons

    Replication and coordination

    Reason about replicas, partitions, quorums, and consensus. Badge: Replica keeper

    1. 0x304Understand replication and consistencySee how replicas improve availability while creating coordination choices. 14 min 14 min
    2. 0x405Apply CAP and quorum reasoningMake the consistency and availability trade-off concrete during a partition. 14 min 14 min
    3. 0x506Coordinate with leaders and consensusLearn what consensus solves and why leader leases need care. 14 min 14 min
  3. Unit 3 · 0/3 lessons

    Scaling and resilience

    Partition workloads, process queued work, and plan recovery. Badge: Systems thinker

    1. 0x607Partition data without creating hotspotsChoose partition keys and recognize skew, fan-out, and rebalancing costs. 14 min 14 min
    2. 0x708Build reliable asynchronous workflowsUse queues to absorb bursts while respecting retries, ordering, and duplicate delivery. 14 min 14 min
    3. 0x809Design for recovery and learn from failuresCombine health signals, graceful degradation, and recovery objectives. 14 min 14 min
Gostou da aula? 😆👍
Apoie nosso trabalho com uma doação: