Design for recovery and learn from failures
Combine health signals, graceful degradation, and recovery objectives.
- Distinguish redundancy from resilience and use SLI, SLO, RTO, and RPO appropriately.
Redundancy provides spare capacity, but resilience also needs failure detection, safe failover, tested recovery, and clear operator signals. An SLI is a measured indicator such as successful requests; an SLO is a target for that indicator over a window. RTO is the target time to restore service after disruption, while RPO is the acceptable amount of data loss measured in time. Graceful degradation preserves essential behavior when optional dependencies fail.
1requests = 10000
2errors = 25
3availability = (requests - errors) / requests
4print(f"success rate: {availability:.2%}")success rate: 99.75%
The example calculates one simple success-rate indicator. A real SLI needs a clear event definition, exclusions, and a measurement window. Pair metrics with logs and traces so teams can find where a request slowed or failed.
Key takeaways
Measure user-visible reliability with clearly defined SLIs and SLOs.
RTO is recovery time; RPO is acceptable data loss.
Resilience includes tested recovery and graceful behavior, not just extra servers.
Lesson quiz
5 questions · pass with 4 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…