Despite the Soft Margin, some problems are intrinsically nonseparable in feature space.
However, based on knowledge of the problem, it may be possible to infer that a nonlinear transformation maps the input feature space
into the feature space
, where the separating hyperplane provides better discrimination between the categories.
The discriminant function in space
is
| (5.43) |
To enable separation, the space is normally of higher dimension than the space
.
This increase in dimensionality would make it computationally expensive to calculate explicitly the coordinates of the samples in space
.
Kernel methods make it possible to work in the transformed space without explicitly calculating its coordinates.
The vector is a linear combination of the training samples (the support vectors in the hard margin case):
| (5.44) |
When evaluating the discriminant function, it is therefore necessary to use the support vectors, at least those associated with a non-negligible parameter .
In effect, an SVM with a kernel identifies some training samples as useful information for determining how close the sample being evaluated is to them.
The bias is calculated directly from equation (5.45) by averaging
| (5.46) |
The most widely used kernels, because they are simple to evaluate, are Gaussian kernels of the form
| (5.47) |
| (5.48) |
The use of kernel functions, together with the possibility of precomputing all combinations
, makes it possible to define a common interface between linear and nonlinear training while effectively maintaining the same level of performance.
It is worth noting that the prediction
takes the form
| (5.49) |
We have therefore seen how a classifier can be constructed by directly imposing a separation geometry between the classes. The hinge-loss formulation, however, also makes it possible to frame SVMs within the more general framework of empirical risk minimization, which will be introduced in the next section.
Paolo medici