CNNs and Image Recognition

How Convolutional Neural Networks see the world—from sliding filters, stride, and pooling to ResNet skip connections and modern Computer Vision.

28 minIntermediateCode Examples

The Core Thesis: If you flatten a high-resolution photo into a 1D list of numbers and feed it into a standard Multilayer Perceptron (MLP), you destroy the 2D spatial layout of the image and create billions of redundant weights. Convolutional Neural Networks (CNNs) solve this by sliding small, reusable 2D filters (kernels) across the image like a flashlight—detecting edges in early layers, textures and parts in middle layers, and complete objects in deep layers.

  1. Beginner Foundations: Why Standard MLPs Fail on Images

To a computer, a color image is a 3D Tensor of pixel brightness values with shape (C,H,W)=(Channels,Height,Width)(C, H, W) = (\text{Channels}, \text{Height}, \text{Width}), where C=3C = 3 for Red, Green, and Blue. Suppose we have a modest 224×224224 \times 224 RGB photo (3×224×224=150,5283 \times 224 \times 224 = 150{,}528 numbers). What goes wrong if we flatten those pixels into a vector and pass them into a standard Dense layer (nn.Linear)?

Problem 1: Parameter Explosion

Connecting 150,528150{,}528 input pixels to just 1,0001{,}000 hidden neurons requires 150,528×1,000=150.5 Million150{,}528 \times 1{,}000 = 150.5\text{ Million} weights in a single layer! The model immediately memorizes the training photos (overfits) and eats gigabytes of VRAM.

1 Dense Layer = 150,528,000 Weights

Problem 2: Zero Translation Invariance

In a Dense layer, every pixel position has its own dedicated weights. If you train an MLP to recognize a cat in the top-left corner, and then shift the exact same cat 1010 pixels to the bottom-right, the MLP fails to recognize it!

Shifted Object = Completely Different Inputs

How CNNs Fix Both Problems at Once (Local Connectivity + Weight Sharing)

Instead of connecting every neuron to all 150,528150{,}528 pixels, a CNN uses a tiny 3×33 \times 3 Filter (Kernel) containing just 99 weights per channel, and slides that exact same filter across every (x,y)(x, y) position of the image! If a filter learns to detect a cat's ear in the top-left, it automatically detects that same ear anywhere in the image (Translation Equivariance) using only a handful of shared weights!

  1. Beginner: How 2D Convolution Works (The Sliding Flashlight)

Imagine holding a 3×33 \times 3 magnifying glass (the Kernel / Filter K\mathbf{K}) over the top-left 3×33 \times 3 patch of an image I\mathbf{I}. At each stop, you multiply the 99 filter weights by the 99 underlying image pixels element-by-element and sum them up (a Dot Product) to produce 11 output pixel in the new Feature Map (Activation Map):

S(i,j)=(I∗K)(i,j)=∑m=0kH−1∑n=0kW−1I(i+m,  j+n)⋅K(m,n)+bS(i, j) = (\mathbf{I} * \mathbf{K})(i, j) = \sum_{m=0}^{k_H - 1} \sum_{n=0}^{k_W - 1} \mathbf{I}(i + m, \; j + n) \cdot \mathbf{K}(m, n) + b

Intuitive Example: How a 3×33 \times 3 Filter Detects Vertical Edges

Suppose we use the classic vertical edge-detector filter K=[1amp;0amp;−11amp;0amp;−11amp;0amp;−1]\mathbf{K} = \begin{bmatrix} 1 & 0 & -1 \\ 1 & 0 & -1 \\ 1 & 0 & -1 \end{bmatrix}. Look at what happens when it slides over two different image patches:

Patch A: Flat Uniform Sky (All 50s)
(50−50)+(50−50)+(50−50)=0(50 - 50) + (50 - 50) + (50 - 50) = 0
Left and right columns cancel out → Output is 0 (No Edge)!
Patch B: Bright Left (100), Dark Right (0)
(100−0)+(100−0)+(100−0)=+300(100 - 0) + (100 - 0) + (100 - 0) = +300
Strong mismatch across boundary → Fires +300 (Vertical Edge Found!)

In classical image processing, engineers designed filters like this by hand. In a CNN, we initialize the filter weights randomly and let Backpropagation learn the exact filter numbers automatically from data!

⚡ Knowledge Check

If an input RGB color image has 33 channels (Cin=3C_{\text{in}} = 3) and we apply a single 3×33 \times 3 convolutional filter to it, how many trainable weights (excluding bias) does that ONE filter have?

A) 27 weights (3 channels × 3 height × 3 width)▼
✓ Correct!Every convolutional filter always spans the full depth CinC_{\text{in}} of its input volume. So a 3×33 \times 3 filter on an RGB image is actually a 3D block of shape (3,3,3)=27(3, 3, 3) = 27 weights that sums across Red, Green, and Blue to output 11 feature map channel!
B) 9 weights (3 height × 3 width)▼
✕ Incorrect.A filter only has 99 weights if the input is a 1-channel grayscale image (Cin=1C_{\text{in}} = 1). On a CinC_{\text{in}}-channel input, each filter has Cin×kH×kWC_{\text{in}} \times k_H \times k_W weights.

  1. Intermediate: Padding, Stride, Pooling & The Output Shape Formula

Three hyperparameters control how a convolutional layer slides across an input of spatial size WinW_{\text{in}} and what output spatial size WoutW_{\text{out}} it produces:

1. Kernel Size (KK)
Usually 3×33 \times 3 or 1×11 \times 1

The spatial window size of the filter. Without padding, a 3×33 \times 3 kernel shrinks the image border by 22 pixels on every layer!

2. Padding (PP)
Usually P=1P = 1 for K=3K = 3

Adds a border of PP zeros around the edges of the image ("Same" Padding) so the output height and width stay the exact same size!

3. Stride (SS)
S=1S = 1 (slide) or S=2S = 2 (skip)

How many pixels the filter jumps on each step. Setting Stride S=2S = 2 skips every other pixel, cutting height and width in half!

The Universal CNN Spatial Output Size Formula
Hout=⌊Hin+2P−KS⌋+1andWout=⌊Win+2P−KS⌋+1H_{\text{out}} = \left\lfloor \dfrac{H_{\text{in}} + 2P - K}{S} \right\rfloor + 1 \qquad \text{and} \qquad W_{\text{out}} = \left\lfloor \dfrac{W_{\text{in}} + 2P - K}{S} \right\rfloor + 1

Pooling Layers: Max Pooling vs. Global Average Pooling

To reduce memory and make the network invariant to small shifts, CNNs periodically downsample the spatial resolution (H,W)(H, W) using Pooling Layers (which contain 00 learnable weights!):

1. Max Pooling (nn.MaxPool2d(2, 2))

Slides a 2×22 \times 2 window with stride 22 and keeps only the single largest number in each 2×22 \times 2 patch. Cuts HH and WW in half (reducing pixels by 75%75\%) while preserving the strongest detected feature!

2. Global Average Pooling (AdaptiveAvgPool2d(1))

Used at the very end of modern CNNs (like ResNet): averages each H×WH \times W feature map down into a single 1×11 \times 1 number per channel, eliminating massive Dense layers and allowing the CNN to accept any input image resolution!

⚡ Knowledge Check

You pass a batch of images with shape (32,3,64,64)(32, 3, 64, 64) into nn.Conv2d(in_channels=3, out_channels=16, kernel_size=3, stride=2, padding=1). What is the exact output tensor shape?

A) (32, 16, 32, 32)▼
✓ Correct!Batch size stays 3232, channels become out_channels = 16, and the spatial formula gives ⌊64+2(1)−32⌋+1=⌊632⌋+1=31+1=32\lfloor \frac{64 + 2(1) - 3}{2} \rfloor + 1 = \lfloor \frac{63}{2} \rfloor + 1 = 31 + 1 = 32.
B) (32, 3, 64, 64)▼
✕ Incorrect.With K=3K=3 and P=1P=1, the spatial size only stays 64×6464 \times 64 when stride=1. Because stride=2, the spatial dimensions are halved to 32×3232 \times 32, and channels change to 1616.

  1. Intermediate to Advanced: Hierarchical Receptive Fields & Classic Architectures

If a 3×33 \times 3 filter only looks at a tiny 3×33 \times 3 patch of pixels, how can a CNN see an entire car or human face? Through the Effective Receptive Field!

A 3×33 \times 3 filter in Layer 1 sees a 3×33 \times 3 patch of raw pixels. A 3×33 \times 3 filter in Layer 2 looks at a 3×33 \times 3 patch of Layer 1 outputs—which covers a 5×55 \times 5 patch of the original image! Stacking three 3×33 \times 3 layers covers a 7×77 \times 7 receptive field using only 3×(32)=273 \times (3^2) = 27 weights per channel instead of 72=497^2 = 49 weights, with three non-linear ReLU folds instead of one!

Landmark ArchitectureYearKey Breakthrough Idea
LeNet-5 (Yann LeCun)1998Pioneered the classic Conv→Pool→Conv→Pool→MLP\text{Conv} \to \text{Pool} \to \text{Conv} \to \text{Pool} \to \text{MLP} blueprint to read handwritten bank checks.
AlexNet2012Combined GPU training, ReLU activations, and Dropout to crush the ImageNet benchmark and ignite the Deep Learning boom.
VGG-16 / VGG-192014Replaced large 7×77 \times 7 filters with uniform stacks of 3×33 \times 3 filters, doubling channels (64→128→256→51264 \to 128 \to 256 \to 512) each time spatial size halved.
ResNet (Residual Networks)2015Introduced Skip (Residual) Connections (y=F(x)+x\mathbf{y} = F(\mathbf{x}) + \mathbf{x}) and Batch Normalization, allowing CNNs to scale past 152152 layers without vanishing gradients!
ConvNeXt & Vision Transformers (ViT)2020–PresentViTs split images into 16×1616 \times 16 patches for Self-Attention[cite: 10], while ConvNeXt modernized ResNets with 7×77 \times 7 depthwise convolutions and GELU.

  1. Advanced: 1×11 \times 1 Convolutions, Depthwise Separable Convs & Modern Vision Tasks

At the Advanced Computer Vision level, three architectural techniques power everything from real-time mobile object detectors (YOLO) to medical image segmentation (U-Net):

1. What is a 1×11 \times 1 Convolution? (Pointwise Bottleneck)

At first glance, a 1×11 \times 1 filter sounds useless because it only looks at 11 pixel at a time! Remember, however, that a filter spans all CinC_{\text{in}} channels: a 1×11 \times 1 convolution mixes information across channels and acts as a channel-dimension bottleneck (e.g., compressing 256256 channels down to 6464 channels before an expensive 3×33 \times 3 conv)!

Used in: ResNet-50 Bottlenecks, Inception, and MobileNet

2. Depthwise Separable Convolutions (9x Faster)

Factors a standard 3D convolution into two cheap steps: (a) a Depthwise Conv (groups=C_in) that applies one spatial 3×33 \times 3 filter per channel independently, followed by (b) a 1×11 \times 1 Pointwise Conv that mixes the channels together—slashing FLOPs by ≈8–9×\approx 8\text{--}9\times!

Powers: MobileNet, EfficientNet, Xception, and ConvNeXt

  1. Beyond Classification: Object Detection (YOLO) & Semantic Segmentation (U-Net)

• Image Classification (ResNet / ViT)[cite: 10]: Outputs 11 class label for the whole image (B,K)(B, K).
• Object Detection (YOLO / Faster R-CNN): Predicts bounding boxes (x,y,w,h)(x, y, w, h) + class labels for multiple objects in real time.
• Semantic Segmentation (U-Net): Uses an Encoder-Decoder CNN with skip connections to classify every individual pixel (B,K,H,W)(B, K, H, W)!

⚡ Knowledge Check

In a Residual Block (Y=F(X)+X\mathbf{Y} = F(\mathbf{X}) + \mathbf{X}), suppose the input tensor X\mathbf{X} has shape (B,64,56,56)(B, 64, 56, 56), and the main branch F(X)F(\mathbf{X}) doubles the channels and halves the spatial size to (B,128,28,28)(B, 128, 28, 28). How do we add X\mathbf{X} to F(X)F(\mathbf{X}) when their shapes do not match?

A) Pass X through a 1x1 Convolution with stride=2 and out_channels=128 on the skip branch▼
✓ Correct!A 1×11 \times 1 projection shortcut with stride=2 and out_channels=128 downsamples X\mathbf{X} spatially to 28×2828 \times 28 while expanding its channels from 6464 to 128128 so F(X)+XprojF(\mathbf{X}) + \mathbf{X}_{\text{proj}} aligns cleanly!
B) Flatten both tensors into 1D vectors and concatenate them▼
✕ Incorrect.Flattening inside a hidden convolutional block destroys the 2D spatial structure needed by downstream convolutional layers.

  1. Visual Explanation: The Hierarchical CNN Feature Pipeline

Look at how an input RGB image is progressively transformed through a modern CNN: spatial dimensions (H,W)(H, W) shrink while feature channels CC expand, moving from low-level edges to high-level object recognition:

  1. Python Implementation: Building a Modern Residual CNN in PyTorch

Here is a complete, runnable PyTorch implementation of a Residual CNN Classifier featuring nn.Conv2d, nn.BatchNorm2d, a Skip Connection, and Global Average Pooling (nn.AdaptiveAvgPool2d(1)):

residual_cnn_vision.pyPython 3.11+ · PyTorch 2.x
import torchimport torch.nn as nnimport torch.nn.functional as F # 1. Define a ResNet-Style Residual Convolutional Blockclass ResidualConvBlock(nn.Module):    def init(self, channels: int):        super().init()        self.conv1 = nn.Conv2d(channels, channels, kernel_size=3, padding=1, bias=False)        self.bn1   = nn.BatchNorm2d(channels)        self.conv2 = nn.Conv2d(channels, channels, kernel_size=3, padding=1, bias=False)        self.bn2   = nn.BatchNorm2d(channels)     def forward(self, x: torch.Tensor) -> torch.Tensor:        identity = x                            # Save input for Skip Connection        out = F.relu(self.bn1(self.conv1(x)))        out = self.bn2(self.conv2(out))        return F.relu(out + identity)           # Residual Addition: F(x) + x! # 2. Assemble the Full Image Recognition CNNclass MiniResNet(nn.Module):    def init(self, numClasses: int = 10):        super().init()        self.stem = nn.Sequential(            nn.Conv2d(3, 32, kernel_size=3, stride=1, padding=1, bias=False),            nn.BatchNorm2d(32),            nn.ReLU()        )        self.resBlock = ResidualConvBlock(32)        self.downsample = nn.MaxPool2d(kernel_size=2, stride=2)        self.globalPool = nn.AdaptiveAvgPool2d((1, 1))  # Pools any HxW down to 1x1!        self.classifier = nn.Linear(32, numClasses)     def forward(self, x: torch.Tensor) -> torch.Tensor:        x = self.stem(x)                        # (B, 3, 32, 32)  -> (B, 32, 32, 32)        x = self.resBlock(x)                    # (B, 32, 32, 32) -> (B, 32, 32, 32)        x = self.downsample(x)                  # (B, 32, 32, 32) -> (B, 32, 16, 16)        x = self.globalPool(x)                  # (B, 32, 16, 16) -> (B, 32, 1, 1)        x = torch.flatten(x, start_dim=1)       # (B, 32, 1, 1)   -> (B, 32)        return self.classifier(x)               # (B, 32)         -> (B, 10) model = MiniResNet(numClasses=10)sampleImages = torch.randn(4, 3, 32, 32)        # Batch of 4 RGB images (32x32)logits = model(sampleImages)totalParams = sum(p.numel() for p in model.parameters())print("Output Logits Shape:", logits.shape, "| Total Params:", totalParams)

Pro Tip (Why We Set bias=False Inside Conv2d Before BatchNorm2d):

Look at nn.Conv2d(..., bias=False) right before nn.BatchNorm2d(32) in the code above! Batch Normalization subtracts the channel mean μ\mu from the convolution output, which completely cancels out any constant bias bb added by the convolution—and BatchNorm already has its own learnable shift parameter β\beta. Setting bias=False saves memory and eliminates redundant parameters!

Key Points

✓CNNs replace massive dense layers with small 2D filters that slide across the image, achieving dramatic parameter savings via Weight Sharing and built-in Translation Equivariance.
✓A convolutional layer mapping CinC_{\text{in}} input channels to CoutC_{\text{out}} output channels with kernel size K×KK \times K contains (Cout×Cin×K×K)+Cout(C_{\text{out}} \times C_{\text{in}} \times K \times K) + C_{\text{out}} trainable parameters.
✓Output spatial size is governed by Wout=⌊Win+2P−KS⌋+1W_{\text{out}} = \lfloor \frac{W_{\text{in}} + 2P - K}{S} \rfloor + 1; setting Padding P=1P = 1 with Kernel K=3K = 3 and Stride S=1S = 1 preserves exact height and width.
✓CNNs learn a natural visual hierarchy: early layers detect simple edges and color gradients, middle layers detect textures and parts, and deep layers recognize complete objects.
✓Residual Connections (y=F(x)+x\mathbf{y} = F(\mathbf{x}) + \mathbf{x}) and Batch Normalization solved vanishing gradients in deep vision networks, enabling ResNets with 50–152+50\text{--}152+ layers.
✓Modern CNNs use 1×11 \times 1 Pointwise Convolutions to mix and compress channels, Depthwise Separable Convolutions for mobile efficiency, and Global Average Pooling to replace heavy fully connected heads.

Common Mistakes

✕ Passing Channels-Last (B, H, W, C) images into PyTorch's nn.Conv2d.

OpenCV, PIL, and Matplotlib store images as (H,W,C)(H, W, C), whereas PyTorch strictly requires Channels-First (B,C,H,W)(B, C, H, W)! Always reorder axes with x.permute(0, 3, 1, 2) or torchvision.transforms.ToTensor().

✕ Forgetting to normalize pixel values from [0, 255] uint8 into float32 before training.

Raw integer pixels in [0,255][0, 255] cause massive initial activations and exploding gradients. Always cast images to float32 scaled to [0,1][0, 1] and normalize by channel mean and standard deviation.

✕ Flattening a large spatial feature map (e.g., 512 x 28 x 28) directly into a huge Linear layer.

Flattening 512×28×28=401,408512 \times 28 \times 28 = 401{,}408 activations into a dense layer createstens of millions of parameters and locks your model to a single input resolution. Use nn.AdaptiveAvgPool2d((1, 1)) first to pool down to 512512 numbers!

✕ Leaving bias=True on Conv2d layers that are immediately followed by BatchNorm2d.

Because BatchNorm subtracts the mean across the batch, any bias added by the preceding convolution is immediately subtracted out and receives zero effective gradient.

The Big Picture

Flattened Dense Network (Blind to 2D Space)

Destroy 2D Grid → 150M+ Redundant Weights → Fails When Objects Shift Position

Convolutional Neural Network (Spatial Inductive Bias)

Slide Shared 3x3 Filters → Edges to Parts to Objects → Global Pool → Recognizes Objects Anywhere!

The important conceptual shift is understanding Inductive Bias: a CNN succeeds on images not because it has more weights than an MLP, but because it has fewer, smarter weights constrained to respect the 2D geometry of the visual world—where nearby pixels form local edges, and an object is the same object no matter where it appears in the frame.

Remember: Whenever your data lives on a spatial or temporal grid (images, video, spectrograms, or 3D medical scans), Convolutions and Residual Connections are the foundational building blocks that turn raw pixels into visual understanding.