Using the Bayesian approach, it would be possible to construct an optimal classifier if both the prior probabilities and the class-conditional densities
were known perfectly.
Normally, this information is rarely available, and the adopted approach is to construct a classifier from a set of examples (training set).
To model , a parametric approach is normally used and, whenever possible, the distribution is assumed to be Gaussian or represented by spline functions.
The most commonly used estimation techniques are Maximum Likelihood (ML) and Bayesian Estimation, which, although different in principle, produce almost identical results. The Gaussian distribution is normally an appropriate model for most pattern recognition problems.
Let us consider the fairly common case in which the probability distribution of the various classes is multivariate Gaussian, with mean
and covariance matrix
.
The optimal Bayesian classifier is