The Bayesian classifier

Using the Bayesian approach, it would be possible to construct an optimal classifier if both the prior probabilities $p(y_i)$ and the class-conditional densities $p(x\vert y_i)$ were known exactly. Normally, such information is rarely available, and the adopted approach is to construct a classifier from a set of examples (training set).

To model $p(x\vert y_i)$, a parametric approach is normally used and, whenever possible, this distribution is assumed to be Gaussian or represented by spline functions.

The most widely used estimation techniques are Maximum Likelihood (ML) and Bayesian Estimation, which, although different in their underlying logic, produce nearly identical results. The Gaussian distribution is normally an appropriate model for most pattern recognition problems.

Let us examine the fairly common case in which the probability distributions of the various classes are multivariate Gaussian distributions with mean $\boldsymbol\mu_i$ and covariance matrix $\boldsymbol\Sigma_i$. The optimal Bayesian classifier is

\begin{displaymath}
\begin{array}{rl}
\hat{y}(\mathbf{x}) = & \argmax_i p(\math...
...\det \boldsymbol\Sigma_i - 2 \log \pi_i \right) \\
\end{array}\end{displaymath} (5.11)

using the negative log-likelihood (Section 2.7). When the prior probabilities $\pi_i$ are equal, equation (5.11) coincides with the problem of finding the minimum Mahalanobis distance (Section 2.4) between the classes in the problem.



Paolo medici
2026-10-01