Convolutional Neural Networks

Figure 5.10: Example of the architecture of a Convolutional Neural Network, LeNet-5.
Image lenet5

Convolutional Neural Networks (CNNs) represent a natural evolution of deep neural networks for processing spatially structured data such as images. Unlike conventional MultiLayer Perceptrons (MLPs), CNNs exploit the local correlation between pixels and spatial redundancy through convolutional layers with shared weights. This architecture makes it possible to learn hierarchical representations that are invariant to translation, while significantly reducing the number of parameters to optimize and improving learning efficiency. The following sections describe the main components of a CNN, including convolutional layers, activation functions, pooling, and typical deep-training architectures.

CNNs are multilayer neural networks similar to MultiLayer Perceptrons, but with a distinctive structure: at least one layer consists of sets of neurons that share weights, known as convolutional layers (convolutional layer).

In a convolutional layer, a neuron's activation depends on the dot product between a kernel (or filter) and a local region of the input:

\begin{displaymath}
a_{i,j} = \sum_{k} \sum_{l} \sum_{m} w_{k,l,m} \, x_{i+k, j+l,m} + b = \mathbf{w}^{\top} \mathbf{x}_{i,j} + b
\end{displaymath} (5.118)

where $\mathbf{w}$ represents the filter weights, $\mathbf{x}_{i,j}$ the local image region centered at $(i,j)$, and $b$ an optional bias.

Convolution can be applied with strides other than 1, thereby producing an activation layer that is subsampled relative to the input. Since filters progressively reduce the dimensions of the activation maps, padding is often introduced to maintain a constant output size.

The activation maps (activation maps) are transformed by a nonlinear activation function, typically a ReLU. A pooling layer is often inserted afterwards to reduce dimensionality and introduce local invariance: max pooling is the most common choice, whereas average pooling and L2-norm pooling were more widely used in the past.

CNNs are designed to exploit multichannel two-dimensional inputs. A CNN generally takes a third- or fourth-order tensor as input; for example, an image $w \times h$ with 3 channels (R,G,B) is a third-order tensor. Each convolutional layer may contain $k$ different kernels, producing activation maps of size $w \times h \times k$ as output.

The architecture of a deep network (Deep Neural Network, DNN) may include:

At the final stage (loss layer), a traditional classification method (MLP, AdaBoost, or SVM) is applied to the reduced representation of the input, while retaining most of the useful information.

DNNs are typically trained using variants of stochastic gradient descent (see section 4.3.3), iteratively optimizing the weights of the filters and fully connected layers on the basis of the classification error.

Paolo medici
2026-10-06