Lesson 2 of 6

Why nonlinearity matters

Without activation functions, a deep network is no more powerful than a single layer. Watch layers stretch the plane and activations bend it until a straight line can separate the data.

Intermediate15 min

In this lesson you will

  • Show that stacked linear layers collapse into one linear layer
  • Describe what linear layers and activation functions do to space
  • Explain why the last layer of a classifier only needs to draw a straight line

The previous lesson wrote a layer as Wx+bW\mathbf{x} + \mathbf{b}. Here is a puzzle. If one layer is good, two layers should be better, and ten better still. But try stacking two layers without anything between them:

W2(W1x+b1)+b2=(W2W1) x+(W2b1+b2).W_2(W_1\mathbf{x} + \mathbf{b}_1) + \mathbf{b}_2 = (W_2 W_1)\,\mathbf{x} + (W_2\mathbf{b}_1 + \mathbf{b}_2).

The right side is just another single layer, with weight matrix W2W1W_2W_1 and bias W2b1+b2W_2\mathbf{b}_1 + \mathbf{b}_2. However many linear layers you stack, they collapse into one. The thing that stops the collapse is the Activation functionThe function a neuron applies to its weighted sum, such as the sigmoid, tanh, or ReLU. Without it, stacking layers of neurons would add no power.Open in glossary between layers. This lesson shows what it does, as geometry.

What a linear layer does to space

Think of a layer with two inputs and two outputs as a machine that moves every point of the plane somewhere new. A Linear transformationA map that can only rotate, stretch, shear, or reflect space, so straight lines stay straight and parallel lines stay parallel. Adding a bias also shifts it (an affine map).Open in glossary can rotate the plane, stretch or squash it along some directions, shear it, or flip it, and the bias then slides everything over. What it can never do is bend: straight lines stay straight, parallel lines stay parallel, and evenly spaced gridlines stay evenly spaced.

That is a problem for data like XOR, where opposite corners belong to the same class. No rotation, stretch, or shear can move the points so a single straight line divides them.

What an activation does

An activation function is applied to each coordinate separately, and it is not linear, so it bends the grid.

  • TanhThe hyperbolic tangent, an S-shaped activation function that squashes any number into the range from -1 to 1. It is a rescaled sigmoid centered on zero.Open in glossary squashes every coordinate into the range from −1 to 1. Points far from the origin get pulled in hard, points near it barely move, so the plane gets curved into a square.
  • ReLUThe rectified linear unit, max(0, z): it passes positive inputs through unchanged and turns negative inputs into zero. The most common activation in deep networks.Open in glossary, max⁡(0,z)\max(0, z), leaves positive coordinates alone and sets negative ones to zero. Whole regions of the plane get folded flat onto an axis.

Alternate the two, linear then activation, linear then activation, and the network can stretch, bend, and fold the plane until the classes come apart.

Watch a network reshape space

A network with two neurons per layer, so every step is a 2D picture. Gridlines show what happens to the whole plane.

Input
Data
Activation
ShowingInput
Training accuracy100%

Try this

  • With XOR and Tanh, press Play and watch the gridlines. During the linear stages they stay straight and parallel. During the activation stages they curve.
  • Switch the activation to None and play again. The grid only ever tilts and stretches, and at the end no line separates the colors. Accuracy is stuck near a coin flip.
  • Try ReLU. Look for places where part of the plane is pressed flat against an axis: that is ReLU zeroing out a coordinate.
  • Try Circle and drag the stage slider slowly. The activation stages squeeze and bend the plane until, at the last stage, a single line separates almost every training point.

The last layer only draws a line

Look at the final stage of the demo. The network’s output layer is linear, followed by a softmax that turns its two numbers into probabilities. A linear function can only split its input space with a straight line (in higher dimensions, a flat plane). So whatever the data looked like at the start, by the last hidden layer it had better be arranged so a straight line works.

This is a useful way to read any classifier: hidden layers learn a new representation of the data in which the answer is easy, and the final layer reads the answer off with a straight cut. In deep networks those intermediate representations are called features, and learning good features automatically is much of what deep learning is about.

A short field guide to activations

ActivationFormulaOutput rangeWhere you see it
Sigmoid11+e−z\frac{1}{1 + e^{-z}}0 to 1Output probabilities for yes-or-no questions; early networks
Tanhtanh⁡z\tanh z−1 to 1Small networks, recurrent networks
ReLUmax⁡(0,z)\max(0, z)0 and upThe default in most deep networks since the early 2010s
GELU, SwiGLUsmooth or gated variants of ReLUGELU: about −0.17 and up; SwiGLU: any valueTransformers, including many large language models

ReLU became the default largely for practical reasons. It is cheap to compute, and for any positive input its slope is exactly 1, so gradients pass through it without shrinking. Sigmoid and tanh flatten out for large inputs, where their slope is nearly zero, which slows learning in deep networks. ReLU has its own failure mode: a neuron whose input is negative for every example outputs 0 everywhere and gets no gradient, so it can stop learning altogether.

Why the demo uses such a narrow networkOptional

Every layer in the demo has exactly two neurons, so every intermediate representation is a picture you can look at. Real networks are much wider, which gives them room to untangle data in more dimensions, and they are also much easier to train. Networks this narrow often get stuck depending on their random starting weights, so for each combination of data and activation we trained from 40 random starts and kept the best one. The training itself is real and runs in your browser when you change a setting.

The demo is inspired by Chris Olah’s 2014 essay Neural Networks, Manifolds, and Topology, which explores what transformations like these can and cannot do. One of its results applies here: layers this narrow cannot truly pull a ring apart from its center. Look closely at the Circle network and you will see it cheats, squeezing the center out through a gap in the ring that happens to fall between the training points. A slightly wider network has room to separate them properly.

Key ideas

  • Stacked linear layers collapse into a single linear layer, so depth needs nonlinearity.
  • Linear layers rotate, stretch, shear, flip, and shift space. Straight lines stay straight.
  • Activation functions bend space: tanh squashes it, ReLU folds part of it flat.
  • A classifier’s last layer splits its input with a straight line, so the hidden layers’ job is to rearrange the data until that works.
  • ReLU is the common default because it is cheap and keeps gradients from shrinking for positive inputs.

Check yourself

Pick an answer to see why it is right or wrong. Nothing is graded. Your first answer is saved in this browser so the question can come back for review.

1A network has three layers with no activation functions between them. What can it represent?
2What does ReLU do to its input?
3At the last stage of the demo a dashed straight line separates the classes. Why is a straight line enough there?

Progress is saved in this browser only.

Up nextBackpropagation
Next
Neural networks
  1. 1Layers are matrix multiplications
  2. 2Why nonlinearity matters
  3. 3Backpropagation
  4. 4Training a network
  5. 5Convolutions
  6. 6Reading handwriting

Try "embedding", "softmax", "overfitting", or "backpropagation".