Image Convolution in Neural Networks

In the previous discussion, we explored the fundamentals of convolution. Now, let's examine how convolution operates specifically within the context of image recognition.

Cross-Correlation Operation

Strictly speaking, what we commonly refer to as a convolution layer actually performs a mathematical operation known as cross-correlation, not true convolution. Let's examine the 2D cross-correlation operation:

In 2D cross-correlation, the convolution window begins at the top-left corner of the input tensor and slides from left to right and top to bottom. When the window moves to a new position, the portion of the tensor contained within the window is element-wise multiplied with the convolution kernel tensor. The resulting tensor is then summed to produce a single scalar value, which becomes the output tensor value at that position. In the example above, the four elements of the output tensor are obtained through 2D cross-correlation, resulting in an output with height 2 and width 2 as follows:

[\begin{split}0\times0+1\times1+3\times2+4\times3=19,\ 1\times0+2\times1+4\times2+5\times3=25,\ 3\times0+4\times1+6\times2+7\times3=37,\ 4\times0+5\times1+7\times2+8\times3=43.\end{split} ]Note that the output size is slightly smaller than the input size. This occurs because the convolution kernel's width and height are greater than 1, and the kernel only performs cross-correlation with positions in the image where it fits completely. Therefore, the output size equals the input size \(n_h \times n_w\) minus the kernel size \(k_h \times k_w\), which is:

[(n_h-k_h+1) \times (n_w-k_w+1). ]This is because we need sufficient space to "move" the kernel across the image. Later, we'll see how to maintain the output size by padding zeros around the image boundaries. Next, let's implement this process in a matrix_correlation function that accepts an input tensor X and a kernel tensor K, returning the output tensor Y.

import torch
from torch import nn
from d2l import torch as d2l

def matrix_correlation(X, K):  #@save
    """Calculate 2D cross-correlation operation"""
    kernel_height, kernel_width = K.shape
    output_height = X.shape[0] - kernel_height + 1
    output_width = X.shape[1] - kernel_width + 1
    Y = torch.zeros((output_height, output_width))
    
    for i in range(output_height):
        for j in range(output_width):
            window = X[i:i + kernel_height, j:j + kernel_width]
            Y[i, j] = torch.sum(window * K)
    return Y

Convolution Layer

A convolution layer performs cross-correlation between the input and convolution kernel weights, then adds a scalar bias to produce the output. Therefore, the two trainable parameters in a convolution layer are the kernel weights and the scalar bias. Just as we randomly initialized fully connected layers previously, we also randomly initialize convolution kernel weights when training models based on convolution layers.

Let's implement a 2D convolution layer based on the matrix_correlation function defined above. In the __init__ constructor, we'll declare weight and bias as two model parameters. The forward propagation function will call matrix_correlation and add the bias.

class Convolution2D(nn.Module):
    def __init__(self, kernel_dimensions):
        super().__init__()
        self.kernel_weights = nn.Parameter(torch.rand(kernel_dimensions))
        self.bias_term = nn.Parameter(torch.zeros(1))

    def forward(self, input_tensor):
        return matrix_correlation(input_tensor, self.kernel_weights) + self.bias_term

Learning Convolution Kernels

We know that different convolution kernels can identify different features in images. However, manually designing these kernels is impractical. So, can we learn the convolution kernel that transforms X into Y?

Let's see if we can learn the kernel that transforms X into Y by examining only the "input-output" pairs. First, we'll construct a convolution layer with its kernel initialized as a random tensor. Then, in each iteration, we'll compare Y with the convolution layer's output using squared error, calculate gradients, and update the kernel. For simplicity, we'll use the built-in 2D convolution layer and ignore the bias.

# Create a 2D convolution layer with 1 output channel and kernel shape (1, 2)
conv_layer = nn.Conv2d(1, 1, kernel_size=(1, 2), bias=False)

# This 2D convolution layer uses four-dimensional input and output format (batch size, channels, height, width)
# where both batch size and channel count are 1
input_data = input_data.reshape((1, 1, 6, 8))
target_output = target_output.reshape((1, 1, 6, 7))
learning_rate = 3e-2  # Learning rate

for epoch in range(10):
    predicted_output = conv_layer(input_data)
    # Calculate loss using L2 norm
    loss = (predicted_output - target_output) ** 2
    # First, zero the gradients
    conv_layer.zero_grad()
    # Calculate loss and perform backpropagation to compute gradients
    loss.sum().backward()
    # Update the convolution kernel
    conv_layer.weight.data[:] -= learning_rate * conv_layer.weight.grad
    if (epoch + 1) % 2 == 0:
        print(f'epoch {epoch+1}, loss {loss.sum():.3f}')

Results:

epoch 2, loss 6.422
epoch 4, loss 1.225
epoch 6, loss 0.266
epoch 8, loss 0.070
epoch 10, loss 0.022

After 10 iterations, the error has decreased sufficiently. Now let's examine the weight tensor of the convolution kernel we learned.

conv_layer.weight.data.reshape((1, 2))

tensor([[ 1.0010, -0.9739]])

Feature Maps and Receptive Fields

The output of a convolution layer is sometimes called a feature map because it can be viewed as a transformation that maps an input to spatial dimensions for the next layer. In convolutional neural networks, for any element x in a given layer, its receptive field refers to all elements (from all previous layers) that might influence the computation during forward propagation.

Note that the receptive field can be larger than the actual input size. Let's use Figure 6.2.1 as an example to explain the receptive field: Given a \(2 \times 2\) convolution kernel, the receptive field of the shaded output element with value 19 consists of the four shaded elements in the input. Suppose the previous output was \(Y\) with size \(2\times2\), and we now add another convolution layer that takes \(Y\) as input and outputs a single element \(z\). In this case, the receptive field of \(z\) on \(Y\) includes all four elements of \(Y\), while the receptive field on the input includes all nine original input elements. Therefore, when an element in a feature map needs to detect input features from a broader area, we can construct a deeper network.

Summary

What we call convolution is essentially using mathematical cross-correlation to extract specific features from images. A convolutional neural network can use multiple kernels to extract different features from images, and these kernels are not manually designed but rather learned through the gradient descent and backpropagation mechanisms of neural networks.

Tags: convolutional-neural-networks computer-vision image-processing pytorch

Posted on Sun, 11 Oct 2026 16:09:37 +0000 by jurdygeorge