Deep Neural Networks

In traditional neural networks, the weights are optimized starting from randomly selected initial values. Although simple to implement, this choice causes performance to deteriorate as the network becomes deeper, whereas shallower architectures (with one or two hidden layers) are generally more stable and easier to train.

Historically, training multilayer neural networks (MLPs) using gradient descent encountered two main obstacles:

Consequently, the idea of using very deep networks to model complex problems was long considered impractical.

Starting in 2012, thanks to the availability of large amounts of data (Big Data), the increasing computational power provided by GPUs, and the development of more effective optimization techniques (such as Adam) 4.3.4, deep neural networks experienced a genuine revival.

A fundamental turning point was the victory of AlexNet (KSH12b) in the 2012 ImageNet Large Scale Visual Recognition Challenge (ILSVRC): for the first time, a deep convolutional neural network substantially outperformed traditional approaches, marking the beginning of the modern era of deep learning.

Since then, deep networks have become the de facto standard in numerous areas of machine learning. In particular, the processing of structured data such as images has benefited enormously from convolutional neural networks (CNNs), which exploit the spatial structure of visual signals to learn hierarchical and translation-invariant representations more efficiently.

Paolo medici
2026-10-01