Continuous integration pipelines
Build and test every change automatically with a fast pipeline: stages, fail-fast ordering, the test pyramid, and dealing with flaky tests.
- Design a CI pipeline that builds and tests every change
- Order and parallelize jobs so the pipeline fails fast and finishes fast
- Balance the test pyramid and handle flaky tests
Byte Bakery now merges to main every day. Without automation, that’s terrifying: who checks that each merge still builds and passes the tests? A CI pipeline does - on every push and every pull request:
- check out the code on a clean machine;
- install the dependencies;
- run fast checks: formatting, linting, type checks;
- run the unit tests, then slower integration tests;
- build the deployable artifact (a package or container image) and store it.
The pipeline is defined as code, in the repository, so it’s versioned and reviewed like everything else. Tools include GitHub Actions, GitLab CI, Jenkins, CircleCI, Buildkite and Azure Pipelines - the ideas carry over between them.
1stages: [check, test, build]
2
3lint:
4 stage: check
5 image: python:3.12-slim
6 script:
7 - pip install ruff
8 - ruff check .
9
10unit-tests:
11 stage: test
12 image: python:3.12-slim
13 script:
14 - pip install -r requirements.txt
15 - pytest tests/unit --junitxml=report.xml
16 artifacts:
17 reports:
18 junit: report.xml
19
20build-image:
21 stage: build
22 image: docker:27
23 services: [docker:27-dind]
24 script:
25 - docker build -t registry.example.com/byte-bakery/shop:$CI_COMMIT_SHORT_SHA .
26 - docker push registry.example.com/byte-bakery/shop:$CI_COMMIT_SHORT_SHA
27 rules:
28 - if: $CI_COMMIT_BRANCH == "main"Fast feedback
A pipeline is only useful if people wait for it. Aim for under 10 minutes:
- Fail fast: run the cheapest checks that catch the most problems first. A 20-second lint failure shouldn’t wait behind a 15-minute test suite.
- Parallelize independent jobs, and split big test suites across machines.
- Cache dependencies (next lesson) and only rebuild what changed.
- Stop the line: when
mainis red, fixing it (or reverting) is everyone’s top priority. A pipeline that’s always red teaches people to ignore it.
Try it
Run Byte Bakery’s pipeline
Here is the pipeline as a job graph on a timeline.
- Which chain of jobs decides the total time? That’s the critical path - speeding up anything else won’t help.
- Make unit-tests fail. Which jobs are skipped? What still runs?
- Make deploy-staging fail. When does rollback run?
- ✓ success0-1 min
- ✓ success0-2 min
- ✓ success2-6 min
- ✓ success2-8 min
- ✓ success6-9 min
- ✓ success9-11 min
- ⊘ skipped
- ✓ success11-12 min
Click a job to make it fail (click again to fix it). Bars show when each job runs; jobs without a dependency between them run at the same time.
The test pyramid
Not all tests are equal. The test pyramid suggests many fast unit tests at the bottom, fewer integration tests (with a real database or service) in the middle, and a handful of slow end-to-end tests (a browser clicking through checkout) at the top. An upside-down pyramid - mostly end-to-end tests - gives a slow, flaky pipeline.
A flaky test passes and fails on the same code - because of timing, test order, shared state or the network. Flaky tests are poison: people learn to click “retry” and stop trusting red builds. Detect them (same commit, different results), quarantine them so they don’t block merges, and fix or delete them quickly.
Key takeaways
A CI pipeline builds and tests every change on a clean machine, defined as code in the repo.
Keep it fast: fail fast, parallelize, cache - and the critical path sets the total time.
Many unit tests, fewer integration tests, few end-to-end tests.
Stop the line when main breaks, and quarantine and fix flaky tests.
Lesson quiz
7 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: automate DevOps chores in Python
Write the small Python tools DevOps teams really build - pipeline runners, plan checkers, metric calculators, scanners - and run them against sample inputs. They run locally in your browser; no servers or cloud accounts needed.
Find the critical path
Each input line is a job: name minutes needs... (the jobs it waits for, possibly none). Jobs start as soon as all their needs have finished; independent jobs run in parallel.
Print each job in input order as unit-tests: 2-6 (start and finish minute), then total: 11 min and the critical path, critical path: install -> unit-tests -> build-image -> deploy-staging: start from the job that finishes last and repeatedly step to the need that finished last (the first listed one on ties).
- Byte Bakery pipeline
- Out of order
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Hunt flaky tests
Each input line is a test result from CI: commit test pass|fail, oldest first. Classify every test (alphabetically):
flaky (mixed results on 2 commits) -> quarantineif it both passed and failed on at least one commit (count those commits);- otherwise
failing -> fix nowif every run on the newest commit (the commit of the last line) failed; - otherwise
stable.
Finish with quarantine: test_a, test_b (or quarantine: none).
- A week of CI
- All green
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…