Neural Networks

Ensemble Learning techniques construct complex models by combining several relatively simple classifiers. Neural networks follow a different philosophy: instead of combining independent classifiers, they construct a single parametric model capable of simultaneously learning a representation of the data and the classification function.

Figure 5.8: Example of a neural-network topology.
Image fig_nn

Research in Machine Learning (and Computer Vision in general) has always sought inspiration from the human brain when developing algorithms. Artificial neural networks (artificial neural networks, ANN) are based on the concept of an “artificial neuron,” that is, a structure which, similarly to the neurons of living organisms, applies a nonlinear transformation (called the activation function) to the weighted contributions of the neuron's different inputs:

\begin{displaymath}
y_{k} = f_{k} \left(\sum_i w_{k,i} x_{k,i} + b_{k}\right)
\end{displaymath} (5.107)

where $x_{k,i}$ are the various inputs associated with the $k$-th neuron, with weights $w_{k,i}$, $y_{k}$ is the neuron's response, and the strongly nonlinear activation function $f$ is normally a step function, a sigmoid, or a logistic function. The bias $b$ is sometimes simulated by means of a constant input $x_{k}=+1$.

The simplest neural network, consisting of an input stage and an output stage, is analogous to the perceptron model (perceptron) introduced by Rosenblatt in 1957. Like the brain of living organisms, an artificial neural network consists of interconnected artificial neurons.

The geometry of a feedforward neural network, the topology normally used in practical applications, is that of a MultiLayer Perceptron (MLP) and consists of multiple hidden layers of neurons connecting the input stage to the output stage, which will serve as the input to the next layer. A multilayer perceptron can be viewed as a function

The training phase consists of estimating the weights $w^{k}_i$ that minimize the error between the training labels and the values predicted by the network $f_w(\mathbf{x})$:

\begin{displaymath}
S(\mathbf{w}) = \sum_i \left\Vert \mathbf{y}_i - f_\mathbf{w}(\mathbf{x}_i) \right\Vert^2
\end{displaymath} (5.108)

The weights $w^{k}_i$ can be estimated using standard optimization techniques. Typically, back propagation is used; this is effectively gradient descent combined with the chain-rule for computing derivatives, since MLPs are layered structures.



Subsections
Paolo medici
2026-10-01