An example of dimensionality reduction for classification purposes is Linear Discriminant Analysis (Linear Discriminant Analysis) (Fisher, 1936).
An analysis of how PCA works (Section 2.9.1) shows that this technique merely maximizes information without distinguishing among any classes that may constitute the problem: PCA does not take into account the fact that the data represent different categories. PCA is not a true classifier, but it is a useful technique for simplifying the problem by reducing its dimensionality. LDA, on the other hand, seeks a projection that maximizes the separation between classes relative to the variability within the classes.
For a two-class problem, the best Bayesian classifier is the one that identifies the decision boundary formed by the hypersurface along which the conditional probabilities of the two classes are equal.
If we assume that the two classes in the binary problem have multivariate Gaussian distributions and the same covariance matrix
, it is easy to show that the Bayesian decision boundary, equation (5.11), becomes linear.
LDA therefore assumes homoscedasticity and, under this assumption, seeks a vector that projects the n-dimensional event space into a scalar space while maximizing class separation and allowing the classes to be separated linearly by a boundary of the form
| (5.13) |
Various metrics can be used to determine this separating surface. The term LDA currently encompasses several techniques, of which Fisher's Linear Discriminant Analysis is the most widely used in the literature.
It can be shown that the projection maximizing the separation between the two classes from a “statistical” perspective, that is, the decision hyperplane, is obtained as
| (5.14) |
| (5.15) |
| (5.16) |
This decision boundary coincides with the Bayesian classifier under the assumption that the classes are described by Gaussian distributions with identical covariance matrices. Under these conditions, the decision boundary is linear. However, this approach requires explicitly modeling the statistical distribution of the observations within each class. An alternative is to model directly the posterior probability of the categories without making assumptions about the data distribution: this is the idea underlying logistic regression.
Paolo medici