Classification and machine learning techniques play a predominant role in Artificial Vision. The large amount of information that can be extracted from a video sensor far exceeds that obtainable from other sensors, but exploiting this wealth of information requires complex techniques.
As already stated, statistics, classification, and model fitting can essentially be viewed as different aspects of a single subject. Statistics seeks the most appropriate way, from a Bayesian perspective, to extract the hidden parameters (of the state or model) of a system, possibly affected by noise, attempting, given the inputs, to return the most probable output, whereas classification proposes techniques and methods for modeling the system efficiently. Finally, if the exact model underlying a physical system were known, any classification problem could be reduced to an optimization problem. For these reasons, it is therefore neither easy nor clear to determine where one subject ends and the other begins.
The classification problem can be reduced to determining the parameters of a generic model that allows the problem to be generalized given a limited number of examples.
A classifier can be viewed in two ways, depending on the type of information that the system is intended to provide:
In the first case, a classifier is represented by a generic function
Because of both the infinite number of possible functions and the lack of further specific information about the form of the problem, the function cannot be a precisely specified function, but is represented by a parameterized model of the form
The training phase is based on a set of examples (training set) consisting of pairs
, and through these examples the training phase must determine the parameters
of the function
that minimize, according to a given metric (cost function), the error on the training set itself.
To train the classifier, it is therefore necessary to determine the optimal parameters
that minimize the error in the output space: classification is also an optimization problem.
For this reason, machine learning, model fitting, and statistics are closely related research areas.
The same considerations used for Kalman or Hough, together with everything discussed in the chapter on least-squares model fitting, can be used for classification, while specific classification algorithms can be used, for example, to fit a set of noisy observations to a curve.
It is normally impossible to produce a complete training set: it is not always possible to obtain every type of input-output association so as to systematically map the entire input space into the output space and, even if this were possible, storing the memory required to represent such associations as a Look Up Table would still be costly. These are the main reasons for using models in classification.
The fact that the training set cannot cover all possible input-output combinations, combined with the generation of a model optimized for such incomplete data, can prevent the training process from generalizing: elements not present in the training set might be classified incorrectly because of excessive adaptation to the training set (the overfitting problem). This phenomenon is normally caused by an optimization phase that focuses more on reducing the output error than on generalizing the problem.
Returning to the different ways of viewing a classifier, it is often simpler and more generalizing to derive directly from the input data a surface in that separates the categories in the n-dimensional input space.
A new function
can be defined that associates one and only one output label with each input group, in the form
Expression (5.1) can always be converted into form (5.2) through majority voting:
| (5.3) |
From this perspective, the classifier is a function that directly returns the symbol most similar to the supplied input.
The training set must now associate a single output class
with each input (each element of the space).
This way of viewing a classifier usually reduces computational complexity and resource usage.
If function (5.1) effectively represents a transfer function, or a response,
function (5.2) can be viewed as a partition of space
in which each region, generally very complex and noncontiguous in the input space, is associated with a single class.
For the reasons given above, it is not physically possible to construct an optimal classifier (except for problems of very small dimension or for simple, perfectly known models), but several general-purpose classifiers exist that can be considered suboptimal depending on the problem and the required performance.
In the case of classifiers (5.2), the problem is to obtain an optimal partition of the space; therefore, a set of fast primitives is required, using little memory even for high values of ,
whereas in case (5.1), a function that models the problem very accurately while avoiding specialization is explicitly required.
The information (features) that can be extracted from an image to enable its classification is diverse. In general, directly using the image's grayscale or color values is rarely appropriate in practical applications because these values are normally affected by scene illumination and, above all, because they would represent a very large input space that is difficult to manage. It is therefore necessary to extract essential information (features) from the image region to be classified, describing its appearance as accurately as possible. For this reason, all the theory presented in section 7 is widely used in machine learning. Both Haar features, because of their extraction speed, and Histograms of Oriented Gradients (HOG, Sec. 7.2), because of their accuracy, are widely used. As a compromise and generalization of these two feature classes, Integral Channel Features (ICF, Sec. 7.3) have recently been proposed.
To reduce the complexity of the classification problem, it can be divided into several layers addressed independently: a first layer transforms the input space into the feature space, while a second layer performs the actual classification starting from the feature space.
From this perspective, classification techniques can be divided into three main categories:
Recently, Representation learning techniques constructed from multiple layers arranged in cascade (Deep Learning) have been highly successful in solving complex classification problems.
Among the techniques for transforming the input space into the feature space, PCA, an unsupervised linear technique, is particularly important. Principal Component Analysis (section 2.9.1) is a technique that reduces the number of inputs to the classifier by removing linearly dependent or irrelevant components, thereby reducing the dimensionality of the problem while attempting to preserve as much information as possible.
Regarding models and general-purpose modeling techniques, the following are widely used:
|