Um momento
0x20Lesson 3 of 11

Linear models and gradient descent

Fit a line by minimizing a loss, step by step, and see how the learning rate decides whether training works.

25 min 6-question quiz 2 code exercises
By the end of this lesson you can
  • Write a linear model and its mean squared error loss
  • Explain gradient descent and the role of the learning rate
  • Implement gradient descent with numpy

A linear model predicts with a weighted sum of features plus a bias: with one feature, y^=wx+b\hat{y} = wx + b. Training means finding the ww and bb that make predictions closest to the targets. “Closest” is measured by a loss function; for regression the usual choice is mean squared error:

L(w,b)=1n∑i=1n(wxi+b−yi)2L(w, b) = \frac{1}{n} \sum_{i=1}^{n} \left(w x_i + b - y_i\right)^2

Gradient descent minimizes it by repeatedly stepping downhill. The gradient ∇L\nabla L points uphill, so each step moves the parameters against it, scaled by the learning rate η\eta:

w←w−η∂L∂w,b←b−η∂L∂bw \leftarrow w - \eta \frac{\partial L}{\partial w}, \qquad b \leftarrow b - \eta \frac{\partial L}{\partial b}

This same loop - predict, measure loss, compute gradients, step - trains everything from linear regression to large language models.

Try it

Gradient descent, step by step

Start at w=0,b=0w = 0, b = 0. With learning rate 0.01, take 100 steps a few times and watch the loss fall. Then try 0.1, and finally 0.6. What happens when the steps are too big?

Learning rate:
y = 0.00x + 0.00
Loss (log scale) over 0 steps · now 20.0
gradient.py
1import numpy as np
2x = np.array([0.0, 1.0, 2.0])
3y = np.array([1.0, 3.0, 5.0])
4w, b = 0.0, 0.0
5error = w * x + b - y
6print(2 * np.mean(error * x), 2 * np.mean(error))
Output
-8.666666666666666 -6.0

Key takeaways

  • A linear model predicts y^=wx+b\hat{y} = wx + b; training minimizes a loss like mean squared error.

  • Gradient descent repeatedly steps against the gradient, scaled by the learning rate η\eta.

  • Too large a learning rate diverges; too small crawls. Scaled features help.

Lesson quiz

6 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: write Python

Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

Exercise 1

Mean squared error

+25 XP

Line 1 is w b; line 2 holds x values; line 3 holds y values. Print the mean squared error of y^=wx+b\hat{y} = wx + b with 4 decimals, using numpy.

  • A rough line
  • A perfect line
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Fit a line with gradient descent

+25 XP

Line 1 is LEARNING_RATE STEPS; line 2 x values; line 3 y values. Start at w=b=0w = b = 0 and run gradient descent on mean squared error, using

∂L∂w=2n∑(wxi+b−yi) xi,∂L∂b=2n∑(wxi+b−yi)\frac{\partial L}{\partial w} = \frac{2}{n} \sum (wx_i + b - y_i)\,x_i, \qquad \frac{\partial L}{\partial b} = \frac{2}{n} \sum (wx_i + b - y_i)

Print step S: loss L (4 decimals) after steps 1, 10, 100 and the last step (each once, if ≤ STEPS), then w=W b=B (3 decimals). If the loss ever exceeds 10610^6, print diverged at step S and stop.

  • Converges
  • Learning rate too big
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Gostou da aula? 😆👍
Apoie nosso trabalho com uma doação: