Um momento
0x20Lesson 3 of 12

How machines learn from data

Features and labels, the guess-check-adjust loop, train/test splits, overfitting, and why data quality decides everything.

20 min 6-question quiz 2 code exercises
By the end of this lesson you can
  • Name the features and the label in a dataset
  • Describe the learning loop: predict, measure the error, adjust
  • Explain why models are tested on data they never trained on

Imagine teaching a child what a dog is. You don’t hand them a definition; you point at dogs (and cats, and foxes) and say which is which. After enough examples, they recognize dogs they’ve never seen. Machine learning works the same way.

A dataset is a table of examples. Each one has:

  • Features - the inputs the model gets to look at: size, weight, color, number of rooms…
  • A label - the right answer we want it to predict: “dog”, “spam”, “310,000310{,}000 dollars”.

Learning means finding a rule that maps features to labels well enough that it also works on new examples.

Try it

Feature or label?

A hospital wants to predict whether a patient will be readmitted within 30 days. Sort each column: is it an input the model can use, the answer it should predict, or something that shouldn’t be used at all?

0 of 6 sortedScore 0/0
  • “Patient age”

  • “Number of previous hospital visits”

  • “Readmitted within 30 days (yes/no)”

  • “Length of the current stay in days”

  • “Date of the readmission visit”

  • “Patient’s record ID number”

The learning loop

Almost every learning algorithm runs the same loop, thousands or millions of times:

  1. Predict with the current settings (the model’s parameters).
  2. Compare the predictions with the true labels.
  3. Measure how wrong they were with a single number, the loss.
  4. Adjust the parameters a little, in the direction that lowers the loss.
  5. Repeat until the loss stops improving.

The program below learns one parameter - a price per square meter - starting from a terrible guess. Run it and watch the guess improve.

guess_and_adjust.py
1sizes = [50, 80, 120]          # square meters
2prices = [150000, 240000, 360000]
3
4price_per_m2 = 1000             # a bad first guess
5for round_number in range(1, 6):
6    errors = [price_per_m2 * size - price for size, price in zip(sizes, prices)]
7    average_error = sum(errors) / len(errors)
8    print(f"round {round_number}: guess {price_per_m2:.0f}, average error {average_error:.0f}")
9    price_per_m2 -= average_error / 100   # adjust against the error
Output
round 1: guess 1000, average error -166667
round 2: guess 2667, average error -27778
round 3: guess 2944, average error -4630
round 4: guess 2991, average error -772
round 5: guess 2998, average error -129

Testing on unseen data

A model that scores 100% on the examples it studied might just have memorized them, like a student who memorized last year’s exam answers. To find out whether it really learned, we hide some data from it.

Before training, split the dataset: usually 70-80% for training and the rest for testing. The model never sees the test set while learning, so its score there predicts how it will do in the real world.

Two ways to fail:

  • Underfitting - the model is too simple to capture the pattern. Bad on training and test data.
  • Overfitting - the model memorizes the training data, noise and all. Great on training data, bad on test data.

Try it

Watch overfitting happen

Filled dots are training data; hollow dots are held back for testing. Raise the curve’s complexity (its degree) from 1 upward. At degree 1 the straight line underfits. Keep going: training error keeps falling all the way to degree 9 - but at which degree is the error on the held-out points lowest?

● training · ○ validation
Error by degree (log scale): solid training, dashed validation

Training error

0.378

Validation error

0.277

Underfitting: too simple to follow the curve.

Key takeaways

  • Features are the inputs; the label is the answer to predict.

  • Learning is a loop: predict, compare, measure the loss, adjust, repeat.

  • Always judge a model on a test set it never trained on.

  • Overfitting = memorizing the training data; underfitting = too simple to learn the pattern.

Lesson quiz

6 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: write Python

Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

Exercise 1

Split the data

+25 XP

The first line is the training fraction (like 0.8). Every other line is one example. Keep the order: the first int(count * fraction) examples are training data and the rest are test data. Print train: N, test: M, then each test example on its own line.

  • Ten rows
  • Seven rows
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Diagnose the model

+25 XP

Each line is name train_accuracy test_accuracy (percentages). For each model print name: underfitting if both scores are below 70, name: overfitting if the training score is more than 10 points higher than the test score, and name: good otherwise.

  • Three models
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Gostou da aula? 😆👍
Apoie nosso trabalho com uma doação: