Um momento
0xA0Lesson 11 of 16

State, locking and drift

Understand Terraform state, share it safely with remote backends and locks, and detect and resolve drift between code and reality.

26 min 7-question quiz 2 code exercises
By the end of this lesson you can
  • Explain what Terraform state is for and how to store and protect it
  • Explain why state locking matters when several people or pipelines apply
  • Detect drift and decide whether to revert it, codify it or ignore it

How does Terraform know that aws_instance.web in your file is the server i-0a1b2c3d in the cloud? Through its state: a JSON file mapping each resource address to the real object’s ID and last known attributes. Plans compare three things - the configuration, the state, and (after a refresh) the real world.

Rules for state:

  • Store it in a remote backend (an S3 bucket, Azure Storage, Google Cloud Storage, HCP Terraform...) so the whole team and CI share one copy - not on a laptop.
  • Lock it during operations, so two applies can’t run at once and corrupt it.
  • Treat it as sensitive: it can contain passwords and keys. Encrypt it and restrict access.
  • Never edit it by hand. Use terraform import to adopt existing resources, terraform state mv (or moved blocks) when renaming, and terraform state rm to stop managing something.
backend.tf
1terraform {
2  backend "s3" {
3    bucket       = "byte-bakery-terraform-state"
4    key          = "staging/web.tfstate"
5    region       = "eu-west-1"
6    encrypt      = true
7    use_lockfile = true   # take a lock in the bucket while plan or apply runs
8  }
9}

Drift

Drift is any difference between the code and reality, usually from changes made outside Terraform: someone resized a server in the console during an incident, or opened a port “just for a minute”. Detect it by running terraform plan -refresh-only (or a normal plan) on a schedule, and resolve each difference deliberately:

  • Revert it - apply the code again - when the manual change was a mistake or a security risk.
  • Codify it - update the code to match - when the change should stay.
  • Ignore it with lifecycle { ignore_changes = [...] } when something else legitimately manages that attribute, like an autoscaler.

Try it

What should happen to this drift?

Byte Bakery’s nightly drift check found these differences. Decide how to resolve each.

0 of 5 sortedScore 0/0
  • “Someone opened SSH (port 22) to the whole internet while debugging”

  • “During an incident, on-call doubled the database size, and the team agreed to keep it”

  • “The autoscaler keeps changing the group’s desired capacity”

  • “Encryption was switched off on the orders bucket”

  • “The finance team’s cost tool adds a `cost-center` tag to everything”

drift.py
1state = {"instance_type": "t3.small", "port_22_open": "false", "desired_capacity": "2"}
2real = {"instance_type": "t3.large", "port_22_open": "false", "desired_capacity": "5"}
3ignored = {"desired_capacity"}
4
5for key in state:
6    if key not in ignored and state[key] != real[key]:
7        print(f"drift: {key} is {real[key]}, code says {state[key]}")
Output
drift: instance_type is t3.large, code says t3.small

Key takeaways

  • State maps configuration to real resources; keep it remote, locked, encrypted and access-controlled.

  • Never hand-edit state: use import, moved blocks and state commands.

  • Check for drift on a schedule, and resolve it by reverting, codifying or ignoring.

  • Fewer manual changes means less drift: make the code the easy path.

Lesson quiz

7 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: automate DevOps chores in Python

Write the small Python tools DevOps teams really build - pipeline runners, plan checkers, metric calculators, scanners - and run them against sample inputs. They run locally in your browser; no servers or cloud accounts needed.

Exercise 1

Simulate state locking

+25 XP

The first input line is the lock timeout in minutes. Each following line is an apply request, in time order: start_minute person duration. Only one apply can hold the lock. If the lock is free at the request’s start, the apply runs right away; if it will be free within the timeout, it waits and runs as soon as the earlier applies finish (in order); otherwise it fails.

Print alice: applied 0-5, bob: waited 2 min, applied 5-9, or carol: lock timeout after 3 min (held by bob) - the holder being whoever holds the lock when the wait runs out.

  • Busy morning
  • Patient queue
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Report drift

+25 XP

The input has three sections separated by ---: the state (what Terraform last recorded), the real cloud resources, and ignore rules. State and real lines are address key=value ...; ignore lines are address key.

For each resource in state order print aws_instance.web: drifted (instance_type t3.small -> t3.large) (state value then real value, differing attributes in state order, comma-separated, ignored ones skipped), aws_s3_bucket.logs: deleted outside Terraform, or aws_iam_role.app: in sync. Then for real resources not in state, in real order: aws_instance.debug: not managed by Terraform (import or delete it). Finish with drifted: 1, deleted: 1, unmanaged: 1.

  • Nightly check
  • All clean
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Gostou da aula? 😆👍
Apoie nosso trabalho com uma doação: