Neurons and activation functions
Build an artificial neuron, explore activation functions, and see why nonlinearity is essential.
- Compute a neuron’s output as an activation of a weighted sum plus bias
- Compare sigmoid, tanh, ReLU, leaky ReLU and GELU, including their slopes
- Explain why stacked layers need nonlinear activations, and compute a stable softmax
An artificial neuron does two things:
- a weighted sum of its inputs plus a bias:
- an activation function applied to that sum:
The weights say how much each input matters (and in which direction); the bias shifts the threshold. A layer is many neurons sharing the same inputs, so its weights form a matrix and the whole layer is one x @ W + b.
Activation functions
| Activation | Formula | Range | Notes |
|---|---|---|---|
| Sigmoid | (0, 1) | squashes to a probability; flat tails | |
| Tanh | (−1, 1) | zero-centered sigmoid; still flat tails | |
| ReLU | [0, ∞) | the default for hidden layers: cheap, gradient 1 when active | |
| Leaky ReLU | (−∞, ∞) | keeps a small gradient for negative inputs | |
| GELU | about [−0.17, ∞) | a smooth ReLU used in transformers |
What matters most for learning is the slope (derivative): during training, gradients are multiplied by it at every layer. Where the slope is near zero, learning stalls.
Try it
Activation explorer
Pick an activation and slide the input. The dashed curve is its slope.
- Slide sigmoid to x = 5: the output is nearly 1, but the slope is almost 0. Stack ten such layers and the gradient all but vanishes.
- Compare ReLU: its slope is exactly 1 for any positive input - which is why it made deep networks trainable. But for negative inputs its slope is 0 (a “dead” neuron); leaky ReLU fixes that.
σ(x) = 1 / (1 + e^-x)
Solid: the function. Dashed: its derivative (the slope backprop multiplies by).
- output
- 0.731
- slope
- 0.197
1import numpy as np
2
3z = np.array([-3.0, -0.5, 0.0, 0.5, 3.0])
4sigmoid = 1 / (1 + np.exp(-z))
5relu = np.maximum(0, z)
6print("sigmoid:", np.round(sigmoid, 3))
7print("slope: ", np.round(sigmoid * (1 - sigmoid), 3))
8print("tanh: ", np.round(np.tanh(z), 3))
9print("relu: ", relu)
10print("leaky: ", np.where(z > 0, z, 0.1 * z))sigmoid: [0.047 0.378 0.5 0.622 0.953] slope: [0.045 0.235 0.25 0.235 0.045] tanh: [-0.995 -0.462 0. 0.462 0.995] relu: [0. 0. 0. 0.5 3. ] leaky: [-0.3 -0.05 0. 0.5 3. ]
Why nonlinearity is essential
Without activations, a stack of layers is pointless: two linear layers, (x @ W1) @ W2, equal one linear layer with weights W1 @ W2. A hundred linear layers still only draw straight lines. The nonlinearity between layers is what lets depth add power:
1import numpy as np
2
3rng = np.random.default_rng(0)
4x = rng.normal(size=(4, 3))
5W1, W2 = rng.normal(size=(3, 5)), rng.normal(size=(5, 2))
6two_layers = (x @ W1) @ W2
7one_layer = x @ (W1 @ W2)
8print(np.allclose(two_layers, one_layer))
9with_relu = np.maximum(0, x @ W1) @ W2
10print(np.allclose(with_relu, one_layer))True False
Softmax: scores into probabilities
For classification, the last layer outputs one raw score per class - the logits. Softmax turns them into probabilities that are positive and sum to 1:
Computed naively, np.exp(1000) overflows to infinity. Since softmax doesn’t change when you subtract the same number from every logit, subtract the maximum first - the standard stable softmax.
Key takeaways
A neuron computes f(w·x + b); a layer does it for many neurons at once with x @ W + b.
ReLU is the default hidden activation; sigmoid and tanh have flat tails that shrink gradients.
Without nonlinear activations, any stack of layers collapses into a single linear layer.
Softmax turns logits into probabilities; subtract the max first for numerical stability.
Lesson quiz
7 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Activation table
Each input line is a number z. Print a row with z and its sigmoid, tanh, ReLU and leaky ReLU (slope 0.1) values, all to 3 decimals, separated by spaces. Write each activation as a function that works on numpy arrays.
- Five inputs
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
A softmax that never overflows
Each input line holds the logits for one example. Print the softmax probabilities to 3 decimals, then sum=1.000. Logits can be huge (like 1000) - your softmax must subtract each row’s maximum so nothing overflows.
- Huge logits
- Ordinary logits
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…