Despite the Soft Margin, some problems are intrinsically nonseparable in feature space.
However, based on knowledge of the problem, it may be possible to infer that a nonlinear transformation maps the input feature space
into the feature space
, where the separating hyperplane provides better discrimination between the categories.
The discriminant function in space
is
| (5.43) |
To enable separation, the space is normally of higher dimension than the space
.
This increase in dimensionality would make it expensive to calculate explicitly the coordinates of the samples in space
.
Kernel methods make it possible to work in the transformed space without explicitly calculating its coordinates.
The vector is a linear combination of the training samples (the support vectors in the hard margin case):
| (5.44) |
When evaluating the discriminant function, it is therefore necessary to use the support vectors (at least those associated with a non-negligible parameter ).
In practice, kernel SVM identifies some samples from the training set as useful information for determining how close the sample being evaluated is to them.
The bias is calculated directly from equation (5.45) by averaging
| (5.46) |
The most widely used kernels, because they are simple to evaluate, are Gaussian kernels of the form
| (5.47) |
| (5.48) |
The use of kernel functions, together with the possibility of precomputing all combinations
, makes it possible to define a common interface between linear and nonlinear training while effectively maintaining the same level of performance.
It should be noted that the prediction
takes the form
| (5.49) |
We have therefore seen how a classifier can be constructed by directly imposing a separation geometry between the classes. The hinge-loss formulation nevertheless allows SVMs to be brought within the more general framework of empirical risk minimization, which will be introduced in the following section.
Paolo medici