Operate VMs in production
Monitor health, capacity, and incidents across VM fleets.
- Interpret key VM signals and trigger safe responses.
Production VM operations depend on observability: metrics (CPU, memory, disk, network), logs, traces, and health checks. Alerts should be actionable and linked to runbooks. Capacity planning tracks growth trends to avoid resource exhaustion. During incidents, teams triage impact first, then stabilize, then root-cause.
1cpu = [55, 60, 95, 92]
2mem = [70, 72, 88, 91]
3if max(cpu) > 90 or max(mem) > 90:
4 print("page on-call")
5else:
6 print("normal")page on-call
Great operations are boring on purpose: dashboards are clear, alert noise is low, and remediation steps are tested. Game days and failure drills improve response confidence before real incidents happen.
Key takeaways
Interpret key VM signals and trigger safe responses.
VMs are powerful when automation and guardrails stay strong.
Observe before scaling; recover before incidents become outages.
Lesson quiz
5 questions · pass with 4 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Flag unhealthy host state
Read CPU and memory percentages as integers. Print critical if either is > 90; otherwise print ok.
- high memory
- healthy values
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…