Bayes' theorem

The definition of conditional probability allows us to immediately obtain the following fundamental

Theorem 3 (Bayes)   Let $\{\Omega,\mathcal{Y},p\}$ be a probability space. Let the events $y=y_i$ (abbreviated as $y_i$) with $i=1..n$ be a complete system of events of $\Omega$ and $p(y_i)>0 \; \forall i=1..n$.

In this case $\forall y_i \in \mathcal{Y}$ with $p(y_i)>0$ it follows that:

\begin{displaymath}
p(y_i\vert x)=\frac{p(y_i)p(x\vert y_i)}{\sum_{j=1}^n p(y_j)p(x\vert y_j)}
\end{displaymath} (5.7)

and this $\forall i=1..n$.

Bayes' theorem is one of the fundamental elements of the subjectivist, or personal, approach to probability and statistical inference. The system of alternatives $y_i$ with $i=1..n$ is often interpreted as a set of causes; given the initial probabilities of the different causes, Bayes' theorem makes it possible to assign probabilities to the causes given an effect $x$. The probabilities $p(y_i)$ with $i=1..n$ can be interpreted as the a priori knowledge (usually denoted by $\pi_i$), that is, the knowledge available before conducting a statistical experiment. The probabilities $p(x\vert y_i)$ with $i=1..n$ are interpreted as the likelihood, or information concerning $x$, that can be acquired by conducting a suitable statistical experiment. Bayes' formula therefore suggests a mechanism for learning from experience: combining some a priori knowledge about event $y_i$, given by $p(y_i)$, with the knowledge acquired from a statistical experiment, given by $p(x\vert y_i)$, leads to improved knowledge, given by $p(x_i\vert y)$, of event $x_i$, also called the a posteriori probability after the experiment has been performed.

For example, we may have the probability distribution for the color of apples, as well as that for oranges. Using the notation introduced earlier in the theorem, let $y_{1}$ denote the state in which the fruit is an apple, $y_{2}$ the condition in which the fruit is an orange, and let $x$ be a random variable representing the color of the fruit. With this notation, $p(x\vert y_1)$ represents the density function for the color event $x$ conditional on the state being an apple, and $p(x\vert y_2)$ conditional on it being an orange.

During training, it is possible to construct the probability distribution of $p(x\vert y_i)$ for $i$ apple or orange. In addition to this information, the prior probabilities $p(y_{1})$ and $p(y_{2})$ are always known; they simply represent the total number of apples and oranges, respectively.

What we seek is a formula giving the probability that a fruit is an apple or an orange, given that a certain color $x$ has been observed.

Bayes' formula (5.7) provides precisely this:

\begin{displaymath}
p(y_i\vert x) = \frac{p(x\vert y_i)p(y_i)}{p(x)}
\end{displaymath} (5.8)

given the prior information, it makes it possible to calculate the posterior probability that the state of the fruit is $y_i$ given the measured feature $x$. Therefore, after observing a certain $x$ on the conveyor belt and calculating $p(y_1\vert x)$ and $p(y_2\vert x)$, one would classify the fruit as an apple if the first value is greater than the second (and vice versa):

\begin{displaymath}
p(y_1\vert x) > p(y_2\vert x)
\end{displaymath}

that is:

\begin{displaymath}
p(x\vert y_1)p(y_1) > p(x\vert y_2)p(y_2)
\end{displaymath}

In general, for $n$ classes, the Bayesian estimator can be defined by a discriminant function:

\begin{displaymath}
f(x) = \hat{y}(x) = \argmax_i p(y_i\vert x) = \argmax_i p(x\vert y_i) \pi_i
\end{displaymath} (5.9)

It is also possible to calculate an index, given the prior knowledge of the problem, indicating how likely this reasoning is to produce errors. The probability of making an error given an observed feature $x$ depends on the maximum value of the $n$ distribution curves at $x$:

\begin{displaymath}
p(error\vert x) = 1 - \max \left[ p(y_1\vert x), p(y_2\vert x), \dots, p(y_n\vert x) \right]
\end{displaymath} (5.10)

Paolo medici
2026-10-06