In the artificial neuron you met a single neuron: multiply each input by a WeightA number that sets how much one input matters to a neuron, and in which direction. Positive weights push the output up, negative weights push it down.Open in glossary, add them up, add a Bias (neuron)A number a neuron adds to its weighted sum before the activation function. It shifts how easily the neuron switches on, regardless of the inputs.Open in glossary, then apply an activation function. Real networks have thousands to billions of these. Nobody computes them one at a time. This lesson shows the bookkeeping trick that every neural network library uses, and once you see it, the shapes in any architecture diagram start to make sense.
One neuron is a dot product
Take a neuron with three inputs . Before its activation, it computes
Collect the inputs into a vector and the weights into . The sum of products is the dot product , so . A large positive means the input lines up well with the neuron’s weights.
A layer stacks those dot products
A LayerA group of neurons that all read the same inputs. A fully connected layer computes Wx + b for a weight matrix W and bias vector b, then applies an activation function.Open in glossary is a set of neurons that all read the same inputs. Give each neuron its own row in a matrix and the whole layer becomes one Matrix multiplicationCombining a matrix with a vector or another matrix by taking dot products of rows with columns. An (m x n) matrix times an n-vector gives an m-vector.Open in glossary:
Here is the number of inputs and the number of neurons. Row of holds neuron ‘s weights, and entry is that row’s dot product with , plus bias . The activation function is then applied to each separately.
A layer is a matrix multiplication
Each output neuron takes a weighted sum of the inputs plus a bias. Stack those weights in rows and you have a matrix.
positive weightnegative weightThicker means larger.
z₁ = (0.7)(1.0) + (−1.4)(0.5) + (−0.1)(−1.0) + 0.2 = 0.30
Try this
- Select each output neuron in turn, or press Cycle rows. The highlighted lines in the diagram and the highlighted row of W are the same numbers drawn two ways.
- Set every input to 0. Every output now equals its bias, because all the products vanish.
- Find a neuron with mostly positive weights and push the inputs that feed it hardest. Which input moves its output the most? The one with the largest weight in that row.
- Press New weights a few times. The shapes never change: 4 × 3 times 3, plus 4, gives 4.
Shapes are the grammar of networks
The rule for a matrix times a vector is that the inner sizes must match: an matrix accepts only inputs and always returns outputs. Chaining layers means chaining shapes. A network that reads a 28 × 28 image (784 pixels) through layers of 64, 32, and 10 neurons uses
and each layer’s output length is the next layer’s input length. That exact network appears in the last lesson of this track, reading your handwriting.
Counting parameters
Every weight and every bias is a ParameterA number inside a model that training adjusts, such as the slope of a line or a weight in a neural network. Large language models have billions of them.Open in glossary that training adjusts. A layer from inputs to neurons has weights and biases. For the handwriting network above:
| Layer | Weights | Biases | Total |
|---|---|---|---|
| 784 → 64 | 50,176 | 64 | 50,240 |
| 64 → 32 | 2,048 | 32 | 2,080 |
| 32 → 10 | 320 | 10 | 330 |
| Network | 52,650 |
Almost all of the parameters sit in the first layer, because it connects every pixel to every neuron. Large language models follow the same arithmetic, just with matrices thousands of rows wide, which is how they reach billions of parameters.
Whole batches at once
Training rarely feeds one example at a time. Stack input vectors as the rows of a matrix , and the layer handles all of them in one multiplication:
where the bias is added to every row. This is the form libraries actually use. PyTorch’s nn.Linear, for example, stores its weight as an (out × in) matrix and computes .
A whole network as one formulaOptional
Stacking layers composes functions. With an activation function applied elementwise, a network with two hidden layers is
The activations between the matrix multiplications are essential. The next lesson shows what goes wrong without them: the whole formula collapses into a single matrix multiplication.
Key ideas
- A neuron’s weighted sum is a dot product; a layer of neurons on inputs is with of shape .
- Each row of is one neuron. Each column holds the weights for one input.
- Inner dimensions must match, so shapes chain from layer to layer.
- A layer has parameters. The widest connections dominate the count.
- Batches turn into one matrix multiplication, , which GPUs compute very fast.