Using the Bayesian approach, it would be possible to construct an optimal classifier if both the prior probabilities and the class-conditional densities
were known exactly.
Normally, such information is rarely available, and the adopted approach is to construct a classifier from a set of examples (training set).
To model , a parametric approach is normally used and, whenever possible, this distribution is assumed to be Gaussian or represented by spline functions.
The most widely used estimation techniques are Maximum Likelihood (ML) and Bayesian Estimation, which, although different in their underlying logic, produce nearly identical results. The Gaussian distribution is normally an appropriate model for most pattern recognition problems.
Let us examine the fairly common case in which the probability distributions of the various classes are multivariate Gaussian distributions with mean
and covariance matrix
.
The optimal Bayesian classifier is