Bayes' theorem

The definition of conditional probability immediately yields the following fundamental

Theorem 3 (Bayes)   Let $\{\Omega,\mathcal{Y},p\}$ be a probability space. Let the events $y=y_i$ (abbreviated as $y_i$) be such that $i=1..n$ is a complete system of events of $\Omega$ and $p(y_i)>0 \; \forall i=1..n$.

In this case, $\forall y_i \in \mathcal{Y}$ with $p(y_i)>0$, we have:

\begin{displaymath}
p(y_i\vert x)=\frac{p(y_i)p(x\vert y_i)}{\sum_{j=1}^n p(y_j)p(x\vert y_j)}
\end{displaymath} (5.7)

and this $\forall i=1..n$.

Bayes' theorem is one of the fundamental elements of the subjectivist, or personal, approach to probability and statistical inference. The system of alternatives $y_i$ with $i=1..n$ is often interpreted as a set of causes, and Bayes' theorem, given the initial probabilities of the different causes, makes it possible to assign probabilities to the causes given an effect $x$. The probabilities $p(y_i)$ with $i=1..n$ can be interpreted as a priori knowledge (usually denoted by $\pi_i$), that is, knowledge available before performing a statistical experiment. The probabilities $p(x\vert y_i)$ with $i=1..n$ are interpreted as the likelihood, or information concerning $x$, that can be obtained by performing a suitable statistical experiment. Bayes' formula therefore suggests a mechanism for learning from experience: combining some a priori knowledge about the event $y_i$ given by $p(y_i)$ with the knowledge acquired from a statistical experiment given by $p(x\vert y_i)$ yields improved knowledge, given by $p(x_i\vert y)$, of the event $x_i$, also called the a posteriori probability after the experiment has been performed.

For example, we may have the probability distribution for the color of apples, as well as that for oranges. Using the notation introduced earlier in the theorem, let $y_{1}$ denote the state in which the fruit is an apple, $y_{2}$ the condition in which the fruit is an orange, and let $x$ be a random variable representing the color of the fruit. With this notation, $p(x\vert y_1)$ represents the density function for the color event $x$ conditional on the state being an apple, $p(x\vert y_2)$, or an orange.

During training, it is possible to construct the probability distribution of $p(x\vert y_i)$ for $i$ an apple or an orange. In addition to this information, the a priori probabilities $p(y_{1})$ and $p(y_{2})$ are always known; they simply represent the total number of apples relative to the number of oranges.

What we seek is a formula that specifies the probability that a fruit is an apple or an orange, given that a certain color $x$ has been observed.

Bayes' formula (5.7) does precisely this:

\begin{displaymath}
p(y_i\vert x) = \frac{p(x\vert y_i)p(y_i)}{p(x)}
\end{displaymath} (5.8)

given the prior knowledge, it makes it possible to calculate the posterior probability that the state of the fruit is $y_i$ given the measured feature $x$. Therefore, after observing a certain $x$ on the conveyor belt and calculating $p(y_1\vert x)$ and $p(y_2\vert x)$, one will tend to decide that the fruit is an apple if the first value is greater than the second (or vice versa):

\begin{displaymath}
p(y_1\vert x) > p(y_2\vert x)
\end{displaymath}

that is:

\begin{displaymath}
p(x\vert y_1)p(y_1) > p(x\vert y_2)p(y_2)
\end{displaymath}

In general, for $n$ classes, the Bayesian estimator can be defined through a discriminant function:

\begin{displaymath}
f(x) = \hat{y}(x) = \argmax_i p(y_i\vert x) = \argmax_i p(x\vert y_i) \pi_i
\end{displaymath} (5.9)

It is also possible to calculate an index, given prior knowledge of the problem, indicating how prone this reasoning is to errors. The probability of making an error given an observed feature $x$ depends on the maximum value of the $n$ distribution curves at $x$:

\begin{displaymath}
p(error\vert x) = 1 - \max \left[ p(y_1\vert x), p(y_2\vert x), \dots, p(y_n\vert x) \right]
\end{displaymath} (5.10)

Paolo medici
2026-10-01