The Mahalanobis Distance

A widespread problem is determining how likely it is that an element $\mathbf {x}$ belongs to a probability distribution, thereby obtaining an approximate estimate of whether the element is an inlier, that is, belongs to the distribution, or an outlier, that is, lies outside it.

The Mahalanobis distance (Mah36) measures an observation normalized with respect to its variance and is therefore also referred to as the “generalized distance.”

Definizione 8   The Mahalanobis distance of a vector $\mathbf {x}$ with respect to a distribution with mean $\mathbf{\mu}$ and covariance matrix $\mathbf{\Sigma}$ is defined as
\begin{displaymath}
d(\mathbf{x}) = \sqrt { (\mathbf{x} - \mathbf{\mu})^{\top} \mathbf{\Sigma} ^{-1} (\mathbf{x} - \mathbf{\mu}) }
\end{displaymath} (2.22)

the generalized distance of the point from the mean.

This distance can be extended (generalized squared interpoint distance) to the case of two vectors $\mathbf {x}$ and $\mathbf{y}$ that are realizations of the same random variable with covariance distribution $\mathbf{\Sigma}$:

\begin{displaymath}
d(\mathbf{x}, \mathbf{y}) = \sqrt { (\mathbf{x} - \mathbf{y})^{\top} \mathbf{\Sigma} ^{-1} (\mathbf{x} - \mathbf{y}) }
\end{displaymath} (2.23)

In the particular case of a diagonal covariance matrix, the normalized Euclidean distance is recovered, whereas when the covariance matrix is exactly the identity matrix, that is, when the components have unit variance and are uncorrelated, the above formulation reduces to the classical Euclidean distance.

The Mahalanobis distance makes it possible to measure distances between samples whose units of measurement are unknown, effectively assigning an automatic scaling factor to the data.



Subsections
Paolo medici
2026-10-01