Measuring delivery: the DORA metrics
Measure speed and stability together with the four key metrics - deployment frequency, lead time, change failure rate and time to restore.
- Define the four DORA metrics and calculate them from a deployment log
- Explain why speed and stability improve together rather than trade off
- Use metrics to find bottlenecks without turning them into targets to game
Byte Bakery’s CEO asks: “Are we getting better at shipping?” Opinions differ. Data would help.
The DORA research program (DevOps Research and Assessment, now part of Google Cloud) surveyed tens of thousands of professionals over many years and found four key metrics that predict software delivery performance - and, remarkably, organizational performance too:
| Metric | Question | Measures |
|---|---|---|
| Deployment frequency | How often do we deploy to production? | speed |
| Lead time for changes | How long from commit to running in production? | speed |
| Change failure rate | What share of deployments cause a failure needing a fix? | stability |
| Time to restore service | When a failure happens, how fast do we recover? | stability |
1deploys = [ # (hours from commit to deploy, caused a failure?, hours to restore)
2 (30, False, 0), (4, False, 0), (6, True, 2), (3, False, 0), (5, False, 0),
3]
4days = 5
5lead_times = sorted(hours for hours, _, _ in deploys)
6failures = [restore for _, failed, restore in deploys if failed]
7print(f"deploys per day: {len(deploys) / days:.1f}")
8print(f"median lead time: {lead_times[len(lead_times) // 2]} h")
9print(f"change failure rate: {len(failures) / len(deploys):.0%}")
10print(f"mean time to restore: {sum(failures) / len(failures):.1f} h")deploys per day: 1.0 median lead time: 5 h change failure rate: 20% mean time to restore: 2.0 h
Reading the numbers:
- Use the median lead time - one change stuck for a month shouldn’t hide that most ship in hours (or the reverse).
- Recent DORA reports describe elite teams as deploying on demand (many times a day), with lead times under a day, few failed changes, and recovery in under an hour; low performers deploy monthly or less, with lead times of months. The exact thresholds shift from year to year.
- Newer reports add a fifth area, reliability - meeting your availability and performance goals.
Key takeaways
The four keys: deployment frequency, lead time for changes, change failure rate, time to restore.
Top performers are faster and more stable at the same time - small batches make both possible.
Use medians for lead time, and measure trends over time.
Metrics guide improvement; turned into targets, they get gamed.
Lesson quiz
7 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: automate DevOps chores in Python
Write the small Python tools DevOps teams really build - pipeline runners, plan checkers, metric calculators, scanners - and run them against sample inputs. They run locally in your browser; no servers or cloud accounts needed.
Calculate the four keys
The first input line is the number of days covered. Each following line is a deployment: id commit_hour deploy_hour ok, or id commit_hour deploy_hour failed restored_hour (hours since the start of the period).
Print deployment frequency: 1.4 per day, lead time for changes (median): 5.0 h, change failure rate: 29% and time to restore (mean): 1.5 h (or time to restore: no failures). Lead time is deploy − commit; time to restore is restored − deploy.
- A good week
- Release Night era
- Smooth sailing
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Classify delivery performance
Each input line is a team: team deploys_per_week lead_time_hours failure_percent restore_hours. Rate each metric with these simplified thresholds:
| Metric | elite | high | medium | otherwise |
|---|---|---|---|---|
| deploys per week | 7 or more | 1 or more | 0.25 or more | low |
| lead time (hours) | under 24 | under 168 | under 720 | low |
| failure % | 5 or less | 10 or less | 15 or less | low |
| restore (hours) | under 1 | under 24 | under 168 | low |
The overall tier is the weakest of the four. Print oven-team: deploys elite, lead time high, failures elite, restore elite -> overall high.
- Byte Bakery teams
- Boundaries
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…