The Bayesian classifier

Using the Bayesian approach, it would be possible to construct an optimal classifier if both the prior probabilities $p(y_i)$ and the class-conditional densities $p(x\vert y_i)$ were known perfectly. Normally, this information is rarely available, and the adopted approach is to construct a classifier from a set of examples (training set).

To model $p(x\vert y_i)$, a parametric approach is normally used and, whenever possible, the distribution is assumed to be Gaussian or represented by spline functions.

The most commonly used estimation techniques are Maximum Likelihood (ML) and Bayesian Estimation, which, although different in principle, produce almost identical results. The Gaussian distribution is normally an appropriate model for most pattern recognition problems.

Let us consider the fairly common case in which the probability distribution of the various classes is multivariate Gaussian, with mean $\boldsymbol\mu_i$ and covariance matrix $\boldsymbol\Sigma_i$. The optimal Bayesian classifier is

\begin{displaymath}
\begin{array}{rl}
\hat{y}(\mathbf{x}) = & \argmax_i p(\math...
...\det \boldsymbol\Sigma_i - 2 \log \pi_i \right) \\
\end{array}\end{displaymath} (5.11)

using the negative log-likelihood (section 2.7). When the prior probabilities $\pi_i$ are equal, equation (5.11) coincides with the problem of finding the minimum Mahalanobis distance (section 2.4) between the classes in the problem.



Paolo medici
2026-10-06