A single neuron draws a straight line and says “this side yes, that side no”. That was enough for deciding on a bike ride. Now picture data where the yes examples sit in a cluster in the middle and the no examples surround them in a ring. No straight line can separate a center from the ring around it. A single neuron is stuck.
The fix is the idea behind every Neural networkA model built from layers of artificial neurons, where each layer's outputs become the next layer's inputs. Training adjusts all of the weights together.Open in glossary: use several neurons, and feed their outputs into another neuron.
Lines, combined
In the network below, the two inputs (how far across and how far up a point is) go into a row of neurons called a . Each hidden neuron draws its own straight line, shown dashed. Their outputs go into one final neuron, which combines them into the network’s answer. The shading shows that answer across the whole plane.
A tiny neural network
Two inputs, a layer of hidden neurons, and an output. Train it and watch the shading: one neuron can only draw a straight line, but several together can bend.
Try this
- Start with 1 hidden neuron on Circle and press Train. The single dashed line wanders around but can never surround the middle. Accuracy stalls near 68%.
- Switch to 3 hidden neurons and train again. Watch the three lines arrange themselves into a triangle around the center, and the shading bend to match.
- Try 8. The boundary gets rounder, because more lines can approximate a curve more closely.
- Load XOR, where opposite corners belong together, and compare 2 hidden neurons with 3.
Each hidden neuron answers one simple question: which side of my line is this point on? On its own that is not much. But the output neuron can combine the answers: “inside all three lines” picks out a triangle. With more hidden neurons, the combined region can take almost any shape.
That is the core trick of neural networks. Simple pieces, each drawing a straight line, combine into something that can follow curves. Stack more layers, and the second layer can combine the first layer’s shapes into even more complex ones. A network with many layers is called deep, which is where the name Deep learningMachine learning with neural networks that have many layers. It powers modern image recognition, speech recognition, translation, and chatbots.Open in glossary comes from.
Training sets every weight at once
Nobody told the network where to put its lines. It started with random weights, so its first guesses were poor. Each training step measured the loss over all the points and nudged every weight and bias in the direction that lowered it, all at the same time. Over a few hundred steps, the lines drift into useful positions.
Working out the right nudge for every weight, including weights in the hidden layer whose effect on the answer is indirect, is done by a method called backpropagation. It is how every neural network is trained, from this one to the largest language models, and the neural networks track walks through it step by step.
One thing you may have noticed with XOR: two hidden neurons are enough to solve it in principle, yet with two this network gets stuck around 78%. With three or more, it solves the problem. Training is a search, and it can get stuck. Giving a network more neurons than it strictly needs often makes the search easier, and that turns out to be true of very large networks too.
Counting the numbers a network learnsOptional
Each hidden neuron has two weights (one per input) and a bias: 3 numbers. The output side of this demo has 2 output numbers, one per class, each with a weight from every hidden neuron plus a bias; a softmax turns them into probabilities. (That is mathematically the same as one sigmoid output neuron, just written the way larger networks do it.)
With hidden neurons that makes numbers: 7 for one hidden neuron, 17 for three, 42 for eight. The hidden neurons use the tanh activation, an S-shaped curve like the sigmoid but ranging from to .
For comparison, a modern large language model learns billions of numbers, but each one is adjusted by the same kind of step you just watched.
Key ideas
- A single neuron’s boundary is a straight line, so it cannot separate a center from the ring around it.
- A hidden layer gives several neurons, each drawing its own line; the next neuron combines them into curves and enclosed regions.
- Stacking many layers is what makes a network “deep”.
- Training adjusts every weight in the network at once, using backpropagation to work out each nudge.
- Training is a search that can get stuck; extra neurons often make it easier to find a good solution.