Suppose you feed a photo to the kind of network from the earlier lessons. A modest 1000 × 1000 color photo is 3 million inputs. A first layer of 1,000 neurons, each connected to every input, needs 3 billion weights, just for one layer. Worse, it would have to learn separately that a vertical edge in the top-left corner and a vertical edge in the bottom-right are the same thing.
A ConvolutionSliding a small grid of weights (a kernel) across an image and computing a weighted sum at every position, producing a new image called a feature map.Open in glossary solves both problems with one idea: look at a small patch at a time, and use the same weights for every patch.
A small window that slides
A Kernel (convolution)The small grid of weights, often 3 x 3, that a convolution slides across an image. In a convolutional network the kernel values are learned.Open in glossary is a small grid of weights, typically 3 × 3. Place it over a 3 × 3 patch of the image, multiply each pixel by the weight on top of it, and add the nine products. That sum is one pixel of the output. Slide the kernel one pixel over and repeat, across every row, and you get a whole new image called a Feature mapThe output of one convolution kernel applied across a whole image: a grid showing how strongly that kernel's pattern appears at each location.Open in glossary. Bright spots in the feature map mark places where the patch looked like the kernel.
Convolution: a small window sliding over an image
Each output pixel is the sum of a 3 × 3 patch of input pixels, each multiplied by the matching kernel weight.
Try this
- Drag across the input. The orange square is the kernel’s window, and the arithmetic below shows the nine products and their sum for that position.
- With Vertical edges, look at the square: its left edge comes out one color and its right edge the other, while its top and bottom edges vanish. Compare Horizontal edges.
- Put the window in a flat region and check that the sum is zero. Each edge kernel’s weights add up to zero, so it ignores flat areas.
- Turn on Apply ReLU. Only the positive responses survive, which is what a convolutional network passes to its next layer.
- Choose Draw your own, sketch a few strokes, and press Scan the image to watch the output fill in one window at a time.
Why this is the right shape for images
Two properties make convolutions fit images.
- Locality. Each output looks only at a small neighborhood. Nearby pixels are strongly related; distant pixels much less so, at least at first.
- Weight sharing. The same kernel is used at every position, so a pattern is detected the same way wherever it appears. Shift the input and the feature map shifts with it.
Together they slash the parameter count. A 3 × 3 kernel on a grayscale image has 9 weights and a bias, whether the image is 48 pixels wide or 4,000.
From hand-made kernels to learned ones
The kernels in the demo were designed by hand. The Sobel kernels, for example, are a classic edge detector from image processing. In a Convolutional neural networkA neural network whose early layers are convolutions, so the same small learned patterns are detected everywhere in an image. The standard design for image recognition for over a decade.Open in glossary, nobody designs them: the kernel weights are parameters, trained by backpropagation like any other weights.
What the network learns tends to look familiar. When researchers visualized the first-layer kernels of AlexNet, the network that won the 2012 ImageNet image recognition challenge by a wide margin, many looked like edge detectors at various angles, along with color blobs.
Real layers use many kernels at once. A layer with 64 kernels produces 64 feature maps, called channels, and each kernel in the next layer spans all 64 channels, so it combines evidence across them: an edge here and an edge there might make a corner. Stack layers and the patterns grow. Visualizations of trained networks show early layers responding to edges and textures, middle layers to parts such as eyes or wheels, and later layers to whole objects. Between layers, networks usually shrink the feature maps, by sliding the kernel two pixels at a time or by keeping only the largest value in each small block (pooling), so later kernels see a larger region of the original image.
Key ideas
- A convolution slides a small kernel across an image, computing a weighted sum of each patch to build a feature map.
- Locality and weight sharing make it efficient: one small kernel is reused everywhere, regardless of image size.
- Kernels whose weights sum to zero ignore flat regions and respond to change, which is what edge detectors do.
- In a convolutional network the kernels are learned. Early layers find edges; deeper layers combine them into parts and objects.
- A layer’s parameter count depends on kernel size and channel counts, not on the image size.