Lesson 1 of 6

Layers are matrix multiplications

A layer of neurons is one matrix multiplication plus a bias. Seeing it that way explains the shapes in every neural network, and why GPUs train them so fast.

Intermediate14 min

In this lesson you will

  • Write a layer of neurons as z = Wx + b and read off the shape of every piece
  • Connect each row of the weight matrix to one neuron in a network diagram
  • Count the parameters in a layer and in a whole network

In the artificial neuron you met a single neuron: multiply each input by a WeightA number that sets how much one input matters to a neuron, and in which direction. Positive weights push the output up, negative weights push it down.Open in glossary, add them up, add a Bias (neuron)A number a neuron adds to its weighted sum before the activation function. It shifts how easily the neuron switches on, regardless of the inputs.Open in glossary, then apply an activation function. Real networks have thousands to billions of these. Nobody computes them one at a time. This lesson shows the bookkeeping trick that every neural network library uses, and once you see it, the shapes in any architecture diagram start to make sense.

One neuron is a dot product

Take a neuron with three inputs x1,x2,x3x_1, x_2, x_3. Before its activation, it computes

z=w1x1+w2x2+w3x3+b.z = w_1 x_1 + w_2 x_2 + w_3 x_3 + b.

Collect the inputs into a vector x=(x1,x2,x3)\mathbf{x} = (x_1, x_2, x_3) and the weights into w=(w1,w2,w3)\mathbf{w} = (w_1, w_2, w_3). The sum of products is the dot product w⋅x\mathbf{w} \cdot \mathbf{x}, so z=w⋅x+bz = \mathbf{w} \cdot \mathbf{x} + b. A large positive zz means the input lines up well with the neuron’s weights.

A layer stacks those dot products

A LayerA group of neurons that all read the same inputs. A fully connected layer computes Wx + b for a weight matrix W and bias vector b, then applies an activation function.Open in glossary is a set of neurons that all read the same inputs. Give each neuron its own row in a matrix and the whole layer becomes one Matrix multiplicationCombining a matrix with a vector or another matrix by taking dot products of rows with columns. An (m x n) matrix times an n-vector gives an m-vector.Open in glossary:

z=Wx+b,W∈Rm×n, x∈Rn, b,z∈Rm.\mathbf{z} = W\mathbf{x} + \mathbf{b}, \qquad W \in \mathbb{R}^{m \times n},\ \mathbf{x} \in \mathbb{R}^{n},\ \mathbf{b}, \mathbf{z} \in \mathbb{R}^{m}.

Here nn is the number of inputs and mm the number of neurons. Row ii of WW holds neuron ii‘s weights, and entry ziz_i is that row’s dot product with x\mathbf{x}, plus bias bib_i. The activation function is then applied to each ziz_i separately.

A layer is a matrix multiplication

Each output neuron takes a weighted sum of the inputs plus a bias. Stack those weights in rows and you have a matrix.

1.0x₁0.5x₂−1.0x₃0.3z₁−0.8z₂0.5z₃1.4z₄

positive weightnegative weightThicker means larger.

W
×
x
1.00.5−1.0
+
b
0.20.00.30.4
=
z
0.30−0.800.501.35

z₁ = (0.7)(1.0) + (−1.4)(0.5) + (−0.1)(−1.0) + 0.2 = 0.30

1.0
0.5
−1.0
ShapesW is 4 × 3, x has 3, b and z have 4
Neuron 1 output z₁0.30

Try this

  • Select each output neuron in turn, or press Cycle rows. The highlighted lines in the diagram and the highlighted row of W are the same numbers drawn two ways.
  • Set every input to 0. Every output now equals its bias, because all the products vanish.
  • Find a neuron with mostly positive weights and push the inputs that feed it hardest. Which input moves its output the most? The one with the largest weight in that row.
  • Press New weights a few times. The shapes never change: 4 × 3 times 3, plus 4, gives 4.

Shapes are the grammar of networks

The rule for a matrix times a vector is that the inner sizes must match: an (m×n)(m \times n) matrix accepts only nn inputs and always returns mm outputs. Chaining layers means chaining shapes. A network that reads a 28 × 28 image (784 pixels) through layers of 64, 32, and 10 neurons uses

W1∈R64×784,W2∈R32×64,W3∈R10×32,W_1 \in \mathbb{R}^{64 \times 784},\quad W_2 \in \mathbb{R}^{32 \times 64},\quad W_3 \in \mathbb{R}^{10 \times 32},

and each layer’s output length is the next layer’s input length. That exact network appears in the last lesson of this track, reading your handwriting.

Counting parameters

Every weight and every bias is a ParameterA number inside a model that training adjusts, such as the slope of a line or a weight in a neural network. Large language models have billions of them.Open in glossary that training adjusts. A layer from nn inputs to mm neurons has m×nm \times n weights and mm biases. For the handwriting network above:

LayerWeightsBiasesTotal
784 → 6450,1766450,240
64 → 322,048322,080
32 → 1032010330
Network52,650

Almost all of the parameters sit in the first layer, because it connects every pixel to every neuron. Large language models follow the same arithmetic, just with matrices thousands of rows wide, which is how they reach billions of parameters.

Whole batches at once

Training rarely feeds one example at a time. Stack BB input vectors as the rows of a matrix X∈RB×nX \in \mathbb{R}^{B \times n}, and the layer handles all of them in one multiplication:

Z=XW⊤+b,Z∈RB×m,Z = X W^{\top} + \mathbf{b}, \qquad Z \in \mathbb{R}^{B \times m},

where the bias is added to every row. This is the form libraries actually use. PyTorch’s nn.Linear, for example, stores its weight as an (out × in) matrix and computes XW⊤+bXW^{\top} + \mathbf{b}.

A whole network as one formulaOptional

Stacking layers composes functions. With an activation function σ\sigma applied elementwise, a network with two hidden layers is

f(x)=W3 σ ⁣(W2 σ(W1x+b1)+b2)+b3.f(\mathbf{x}) = W_3\, \sigma\!\big(W_2\, \sigma(W_1 \mathbf{x} + \mathbf{b}_1) + \mathbf{b}_2\big) + \mathbf{b}_3.

The activations between the matrix multiplications are essential. The next lesson shows what goes wrong without them: the whole formula collapses into a single matrix multiplication.

Key ideas

  • A neuron’s weighted sum is a dot product; a layer of mm neurons on nn inputs is z=Wx+b\mathbf{z} = W\mathbf{x} + \mathbf{b} with WW of shape m×nm \times n.
  • Each row of WW is one neuron. Each column holds the weights for one input.
  • Inner dimensions must match, so shapes chain from layer to layer.
  • A layer has mn+mmn + m parameters. The widest connections dominate the count.
  • Batches turn into one matrix multiplication, XW⊤+bXW^{\top} + \mathbf{b}, which GPUs compute very fast.

Check yourself

Pick an answer to see why it is right or wrong. Nothing is graded. Your first answer is saved in this browser so the question can come back for review.

1A layer reads 5 inputs and has 3 neurons. What is the shape of its weight matrix W?
2How many learnable numbers does that layer have?
3Why do people write layers as matrix multiplications instead of one neuron at a time?

Progress is saved in this browser only.

Up nextWhy nonlinearity matters
Next
Neural networks
  1. 1Layers are matrix multiplications
  2. 2Why nonlinearity matters
  3. 3Backpropagation
  4. 4Training a network
  5. 5Convolutions
  6. 6Reading handwriting

Try "embedding", "softmax", "overfitting", or "backpropagation".