Neural Networks

Ensemble Learning techniques construct complex models by combining several relatively simple classifiers. Neural networks follow a different philosophy: instead of combining independent classifiers, they construct a single parametric model capable of simultaneously learning a representation of the data and the classification function.

Figure 5.8: Example of a neural-network topology.
Image fig_nn

Research in Machine Learning (and, more generally, Computer Vision) has always sought inspiration from the human brain for the development of algorithms. Artificial neural networks (artificial neural networks, ANN) are based on the concept of an “artificial neuron,” that is, a structure that, similarly to the neurons of living organisms, applies a nonlinear transformation (called the activation function) to the weighted contributions of the neuron's different inputs:

\begin{displaymath}
y_{k} = f_{k} \left(\sum_i w_{k,i} x_{k,i} + b_{k}\right)
\end{displaymath} (5.109)

where $x_{k,i}$ are the various inputs associated with the $k$-th neuron, with corresponding weights $w_{k,i}$, $y_{k}$ is the neuron's output, and the strongly nonlinear activation function $f$ is normally a step, sigmoid, or logistic function. The bias $b$ is sometimes simulated by means of a constant input $x_{k}=+1$.

The simplest neural network, consisting of an input layer and an output layer, is equivalent to the perceptron model (perceptron) introduced by Rosenblatt in 1957. Like the brain of living organisms, an artificial neural network consists of interconnected artificial neurons.

The geometry of a feedforward neural network, the topology normally used in practical applications, is that of a MultiLayer Perceptron (MLP) and consists of multiple hidden layers of neurons connecting the input layer to the output layer, which in turn becomes the input to the next layer.

The training phase consists of estimating the weights $w^{k}_i$ that minimize the error between the training labels and the values predicted by the network $f_w(\mathbf{x})$:

\begin{displaymath}
S(\mathbf{w}) = \sum_i \left\Vert \mathbf{y}_i - f_\mathbf{w}(\mathbf{x}_i) \right\Vert^2
\end{displaymath} (5.110)

The weights $w^{k}_i$ can be estimated using well-known optimization techniques. Normally, the backpropagation technique is used; this is essentially gradient descent with the chain rule for computing derivatives, since MLPs are layered structures.



Subsections
Paolo medici
2026-10-06