Your first model: linear regression
Fit a straight line to data, measure how wrong it is with mean squared error, and grow it to many features.
- Read a line’s slope and intercept as a model’s parameters
- Compute the mean squared error of a line
- Write a multi-feature prediction as a dot product
Suppose you want to predict exam scores from hours studied. Plot the data and the points drift upward in a rough line. Linear regression finds the line that fits best:
Here is the input (hours), (“y-hat”) is the predicted score, is the slope (points gained per extra hour) and is the intercept (the predicted score with zero hours). The two numbers and are the model’s parameters - “training the model” means finding good values for them.
Try it
Fit the line yourself
Move the sliders to make the red lines (the errors between each point and your line) as short as possible overall. Get within 10% of the best possible error, then reveal the best line and compare.
Your sum of squared errors: 1118.0
Measuring “best”: mean squared error
To let a computer find the best line, “best” needs to be a number. For each point, the error (or residual) is the gap between the prediction and the truth. Square each error so negatives don’t cancel positives and big misses count extra, then average:
This mean squared error is the model’s loss. The best line is the one with the smallest MSE.
1hours = [1, 2, 3, 4, 5]
2scores = [52, 55, 61, 60, 68]
3
4def mse(slope, intercept):
5 errors = [slope * x + intercept - y for x, y in zip(hours, scores)]
6 return sum(error ** 2 for error in errors) / len(errors)
7
8for slope, intercept in [(0, 59), (5, 45), (3.7, 48.1)]:
9 print(f"y = {slope}x + {intercept}: MSE = {mse(slope, intercept):.2f}")y = 0x + 59: MSE = 30.20 y = 5x + 45: MSE = 6.80 y = 3.7x + 48.1: MSE = 2.78
More than one feature
Real predictions use many features. A house price might depend on size, bedrooms and age, each with its own weight:
That’s the dot product from the math lesson. Each weight says how much its feature pushes the prediction up (positive) or down (negative).
1weights = [1200, 15000, -800] # per m², per bedroom, per year of age
2bias = 20000
3
4def predict(features):
5 return sum(weight * feature for weight, feature in zip(weights, features)) + bias
6
7print(predict([100, 3, 10])) # 100 m², 3 bedrooms, 10 years old
8print(predict([60, 1, 40]))177000 75000
Key takeaways
Linear regression predicts ; training finds the slope and intercept.
Mean squared error turns “how good is this line” into one number to minimize.
With many features the prediction is a dot product plus a bias: .
Lesson quiz
6 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Score a line
The first line holds a slope and an intercept. Every other line is a point x y. Print each prediction’s error (prediction minus truth) as x=X error=E, then MSE: M with 2 decimal places. Format numbers with :g.
- Study hours
- Perfect fit
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Find the best line exactly
Every input line is a point x y. Compute the least-squares line with these formulas, where and are the averages:
Print slope: M and intercept: B with 2 decimal places, then the prediction for x = 12 as prediction at 12: P (2 decimal places).
- Study hours
- Exact line
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…