Mean and Variance

It is easy to assume that the notion of the mean of a set of numbers is familiar to everyone, at least from a purely intuitive point of view. Nevertheless, this section provides a brief review, gives the relevant definitions, and highlights some interesting aspects.

For $n$ samples of an observed quantity $x$, the sample mean sample mean is denoted by $\bar{x}$ and is given by

\begin{displaymath}
\bar{x} = \frac{1}{n} \sum_{i=1}^{n} x_i
\end{displaymath} (2.1)

By definition, the sample mean is an empirical quantity.

If infinitely many values of $x$ could be sampled, $\bar{x}$ would converge to the theoretical expected value (expected value). This is the law of large numbers (Law of Large Numbers).

The expected value (expectation, mean) of a random variable $X$ is denoted by $\E[X]$ or $\mu$ and can be computed for discrete random variables using

\begin{displaymath}
\E[X] = \mu_x = \sum_{-\infty}^{+\infty} x_i p_X(x_i)
\end{displaymath} (2.2)

and for continuous variables using
\begin{displaymath}
\E[X] = \mu_x = \int_{-\infty}^{+\infty} x p_X(x) dx
\end{displaymath} (2.3)

given knowledge of the probability distribution $p_X(x)$.

We now introduce the concept of the mean of a function of a random variable.

Definizione 5   Let $X$ be a random variable and let $g(x)$ be a measurable function. In the discrete case, the expected value of the random variable $Y=g(X)$ is defined as
\begin{displaymath}
\E[g(X)] =
\sum_i g(x_i) p_X(x_i),
\end{displaymath} (2.4)

whereas in the continuous case it is defined as
\begin{displaymath}
\E[g(X)] =
\int_{-\infty}^{+\infty} g(x) p_X(x)\,dx.
\end{displaymath} (2.5)

There are several functions whose mean has particular significance. When $g(x)=x$, one speaks of first-order statistics (first statistical moment), and, in general, when $g(x)=x^{k}$, one speaks of $k$-order statistics. The expected value is therefore the first-order statistic, while another statistic of particular interest is the second-order moment:

\begin{displaymath}
\E[X^{2}] = \int_{-\infty}^{+\infty} x^{2} p_X(x) dx
\end{displaymath} (2.6)

This statistic is important because it makes it possible to estimate the variance of $X$.

The variance is defined as the expected value of the square of the random variable $X$ after subtracting its expected value, that is, the second-order moment of the function $g(X)= X-\E[X]$:

\begin{displaymath}
\text{var}(X) = \sigma^{2}_X = \E[ (X - \E[X])^{2} ]
\end{displaymath} (2.7)

Since $\E[X]$ is a constant, expanding the square yields the equivalent and widely used form
\begin{displaymath}
\text{var}(X) = \sigma^{2}_X = \E[X^{2}] - \E[X]^{2}
\end{displaymath} (2.8)

The square root of the variance is known as the standard deviation (standard deviation) and has the advantage of having the same unit of measurement as the observed quantity:

\begin{displaymath}
\sigma_X = \sqrt{ \text{var}(X) }
\end{displaymath} (2.9)

We now extend the concepts introduced so far to the multivariate case. The multivariate case can be viewed as an extension to multiple dimensions, where each dimension is associated with a different variable.

The covariance matrix $\Sigma$ is the multidimensional (or multivariable) extension of the concept of variance. It is constructed as

\begin{displaymath}
\Sigma_{ij} =\text{cov}(X_i,X_j)
\end{displaymath} (2.10)

where each element of the matrix contains the covariance between the various components of the random vector $X$. The covariance indicates how the different random variables that make up the vector $X$ are related to one another.

The covariance matrix can be denoted in the following ways:

\begin{displaymath}
\Sigma = \E \left[ (X - \E[X])(X - \E[X])^{\top} \right] = \text{var}(X) = \text{cov}(X) = \text{cov}(X,X)
\end{displaymath} (2.11)

The cross-covariance notation, on the other hand, is unique:

\begin{displaymath}
\text{cov}(X,Y) = \E \left[ (X - \E[X])(Y - \E[Y])^{\top} \right]
\end{displaymath} (2.12)

a generalization of the concept of the covariance matrix. The cross-covariance matrix $\boldsymbol\Sigma$ has, at position $(i,j)$, the covariance between the random variable $X_i$ and the variable $Y_j$:
\begin{displaymath}
\boldsymbol\Sigma = \begin{bmatrix}
\text{cov}(X_1,Y_1) & \...
...cov}(X_1,Y_m) & \cdots & \text{cov}(X_n,Y_m) \\
\end{bmatrix}\end{displaymath} (2.13)

The covariance matrix $\text{cov}(X,X)$ is consequently symmetric.

The covariance matrix describes how the different components are mutually correlated and is also called the scatter matrix (scatter matrix). The inverse of the covariance matrix is called the concentration matrix or precision matrix.

In the scalar case, the correlation coefficient between two random variables $X$ and $Y$ is defined as

\begin{displaymath}
r(X,Y) =
\frac{\text{cov}(X,Y)}
{\sqrt{\text{var}(X)\,\text{var}(Y)}}
=
\frac{\text{cov}(X,Y)}
{\sigma_X \sigma_Y}.
\end{displaymath} (2.14)

The correlation coefficient takes values in the interval $[-1,1]$. Values close to $1$ indicate a strong positive correlation, values close to $-1$ indicate a strong negative correlation, whereas values close to $0$ indicate the absence of linear correlation 2.1.

In the multivariate case, a correlation matrix is defined analogously by normalizing the cross-covariance matrix so that each element represents the correlation coefficient between a component of $X$ and a component of $Y$. In this case as well, all matrix elements belong to the interval $[-1,1]$.



Footnotes

... correlation2.1
A zero correlation coefficient implies the absence of linear correlation but does not necessarily imply statistical independence.
Paolo medici
2026-10-01