Classify with logistic regression
Turn a linear score into a probability with the sigmoid, train with log loss, and pick a decision threshold.
- Explain the sigmoid and why it outputs probabilities
- Compute log loss and understand why it punishes confident mistakes
- Turn probabilities into decisions with a threshold
For classification we want a probability, not an unbounded number. Logistic regression computes a linear score and squashes it with the sigmoid:
which maps any score into : large positive scores approach 1, large negative scores approach 0, and gives exactly 0.5. The result is read as the probability of the positive class.
Training minimizes log loss (binary cross-entropy):
A confident wrong prediction - for an example that is really negative - costs , while a hesitant one costs much less. That pushes the model to be confident only when it’s right.
From probability to decision
A probability isn’t a decision. You choose a threshold: predict positive when . The default 0.5 is rarely special - for a disease screening you may flag patients at because missing a case is worse than an extra test; for blocking a payment you may require .
Try it
Move the threshold
These are a fraud model’s probabilities for 12 transactions. Slide the threshold and watch precision and recall trade off.
| Flagged fraud | Called legitimate | |
|---|---|---|
| Really fraud | 3 true positives | 2 missed (false negatives) |
| Really legitimate | 2 false alarms (false positives) | 5 true negatives |
- 0.95Card used in two countries within an hour (fraud)flagged fraud
- 0.88Large electronics purchase at 3 a.m. (fraud)flagged fraud
- 0.72New device, new shipping address (fraud)flagged fraud
- 0.64Frequent small gift card purchases (legitimate)flagged fraud
- 0.55First purchase at a new store (legitimate)flagged fraud
- 0.48Many failed PIN attempts (fraud)legitimate
- 0.41Travel booking abroad (legitimate)legitimate
- 0.33Small purchase right after a password reset (fraud)legitimate
- 0.18Online order to home address (legitimate)legitimate
- 0.08Fuel at the usual station (legitimate)legitimate
- 0.05Usual grocery store, usual amount (legitimate)legitimate
- 0.02Monthly streaming subscription (legitimate)legitimate
import numpy as np
z = np.array([-4.0, -1.0, 0.0, 1.0, 4.0])
print(np.round(1 / (1 + np.exp(-z)), 3))[0.018 0.269 0.5 0.731 0.982]
Key takeaways
Logistic regression: , a probability between 0 and 1.
Log loss punishes confident mistakes heavily.
Choose the decision threshold from the costs of each kind of error.
Lesson quiz
5 questions · pass with 4 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Compute log loss
Line 1 holds true labels (0 or 1); line 2 the predicted probabilities. Clip probabilities to and print the log loss with 4 decimals.
- Mostly right
- One confident mistake
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Train logistic regression
Line 1 is LEARNING_RATE STEPS. Each following line is x1 x2 label. Start with weights and bias at 0 and run gradient descent on log loss, with :
Print w=[W1, W2] b=B (3 decimals), log loss: L (4 decimals) and training accuracy: A% using a 0.5 threshold.
- Pass or fail from hours studied and slept
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…