All articles
Computer Vision·Architectures

Convolutional Neural Networks, Explained

How convolution, pooling, and stacked layers let a network build up from edges to full objects with far fewer parameters than a fully-connected net.

·2 min read·
2 views
cnncomputer-visiondeep-learning
Notes / Computer Vision

Convolutional Neural Networks, Explained

#cnn#computer-vision#deep-learning

Convolutional Neural Networks, Explained

A fully-connected layer treats every pixel as an independent feature with its own weight — for a modest 224×224 color image, that's over 150,000 inputs before the network has learned anything about what an image is. Convolutional neural networks (CNNs) build in a better prior: nearby pixels are related, and the same pattern (an edge, a curve, a texture) can appear anywhere in the frame.

The Convolution Operation

A convolutional layer slides a small filter — say, 3×3 — across the image, computing a dot product at each position:

(I∗K)(i,j)=∑m∑nI(i+m,j+n) K(m,n)(I * K)(i, j) = \sum_{m} \sum_{n} I(i+m, j+n) \, K(m, n)

II is the image, KK is the filter, and the result is a feature map: high values where the patch under the filter resembles the pattern the filter has learned to detect. Because the same filter is reused at every position, a CNN needs far fewer parameters than a fully-connected layer covering the same image — and a filter that learns to detect a vertical edge works whether that edge appears top-left or bottom-right.

Stacking Layers: From Edges to Objects

Layer depthWhat filters tend to detect
Early (layer 1-2)Edges, colors, simple gradients
MiddleTextures, corners, simple shapes
LateObject parts (eyes, wheels), whole objects

Each layer's output becomes the next layer's input, so filters compose: a middle layer's "wheel-like circle" detector is built from early layers' edge and curve detectors. This hierarchy is why CNNs generalize well to new images — they're recognizing structure, not memorizing pixels.

Pooling: Discarding Position, Keeping Presence

After convolution, a pooling layer (usually max pooling) shrinks each feature map by keeping only the strongest response in each small region:

def max_pool_2x2(feature_map):
    h, w = feature_map.shape
    return feature_map.reshape(h // 2, 2, w // 2, 2).max(axis=(1, 3))

This does two things: it reduces computation for later layers, and it makes the network somewhat tolerant to small translations — if an edge shifts by a pixel or two, the pooled output usually doesn't change.

Receptive Field

A single neuron deep in the network doesn't see the whole image directly — it sees a small patch, whose size (its receptive field) grows with depth as layers compose. Stacking 3×3 convolutions is a deliberate design choice: two stacked 3×3 layers cover the same receptive field as one 5×5 layer, with fewer parameters and an extra nonlinearity in between, which is part of why architectures like VGG and ResNet standardized on small filters stacked deep rather than a few large ones.

Backpropagation, Step by Step

How the chain rule turns a network's forward pass into gradients for every weight, and why the "backward" direction is what makes training tractable.

Newsletter

New articles in your inbox

An email when I publish something new. No spam, unsubscribe anytime.

Double opt-in. See the privacy policy for how your email is handled.

Continue exploring

Browse more technical notes and experiments.

All articles