The Maximum Likelihood Estimator

From a statistical standpoint, the data vector $\mathbf{x} = \left\{ x_1 \ldots x_n \right\}$ consists of realizations of a random variable from an unknown population. The task of data analysis is to identify the population most likely to have generated these samples. In statistics, each population is identified by a corresponding probability distribution, and each probability distribution is associated with a unique parameterization $\boldsymbol\vartheta$: varying these parameters must produce a different probability distribution.

Let $f( \mathbf{x} \vert \boldsymbol\vartheta)$ be the probability density function (PDF) indicating the probability of observing $\mathbf {x}$ given a parameterization $\boldsymbol\vartheta$. If the individual observations $x_i$ are statistically independent of one another, the PDF of $\mathbf {x}$ can be expressed as the product of the individual PDFs:

\begin{displaymath}
f( \mathbf{x} = \left\{ x_1 \ldots x_n \right\} \vert \bold...
...oldsymbol\vartheta) \ldots f_n(x_n \vert \boldsymbol\vartheta)
\end{displaymath} (2.51)

Given a parameterization $\boldsymbol\vartheta$, it is possible to define a specific PDF that expresses the probability of some data occurring rather than other data. In the real-world case, we face the inverse problem: the data have been observed, and we must determine which $\boldsymbol\vartheta$ generated that specific PDF.

Definizione 9   To solve the inverse problem, we define the function $\mathcal{L}: \boldsymbol\vartheta \mapsto [0, \infty)$, the likelihood function (likelihood), as
\begin{displaymath}
\mathcal{L}(\boldsymbol\vartheta \vert \mathbf{x} ) = f (\m...
...rtheta) = \prod_{i=1}^{n} f_i (x_i \vert \boldsymbol\vartheta)
\end{displaymath} (2.52)

in the case of statistically independent observations.

$\mathcal{L}( \boldsymbol\vartheta \vert \mathbf{x} )$ denotes the likelihood of parameter $\boldsymbol\vartheta$ given the observation of events $\mathbf {x}$.

The maximum-likelihood estimator (MLE) principle $\hat{\boldsymbol\vartheta}_{MLE}$, originally developed by R. A. Fisher in the 1920s, selects as the best parameterization the one that provides the best fit of the probability distribution generated by the observed data.

For a Gaussian probability distribution, an additional definition is useful.

Definizione 10   Let $\ell$ be the log-likelihood function (log likelihood), defined as
\begin{displaymath}
\ell = \log \mathcal{L}(\boldsymbol\vartheta \vert x_1 \ldo...
..._n) = \sum_{i=1}^{n} \log f_i (x_i \vert \boldsymbol\vartheta)
\end{displaymath} (2.53)

using the properties of the logarithm.

The best estimate of the model parameters is the one that maximizes the likelihood, or equivalently the log-likelihood

\begin{displaymath}
\hat{\boldsymbol\vartheta}_{ML} = \argmax_{\boldsymbol\vart...
...heta} \sum_{i=1}^{n} \log f_i (x_i \vert \boldsymbol\vartheta)
\end{displaymath} (2.54)

since the logarithm is a monotonically increasing function.

In the literature, the optimum estimator is sometimes defined not as the maximum of the likelihood function, but as the minimum of its negative

\begin{displaymath}
\hat{\boldsymbol\vartheta}_{ML} = \argmin_{\boldsymbol\vart...
...um_{i=1}^{n} \log f_i (x_i \vert \boldsymbol\vartheta) \right)
\end{displaymath} (2.55)

that is, the minimum of the negative log likelihood.

This formulation is particularly useful when the noise distribution is Gaussian. Let $(x_i,y_i)$ be realizations of the random variable. In fact, for a generic function $y_i = g(x_i ; \boldsymbol\vartheta) + \epsilon$ with normally distributed noise, constant variance, and zero mean, the likelihood is

\begin{displaymath}
\mathcal{L}(\boldsymbol\vartheta \vert \mathbf{x} ) = \prod...
...- g(x_i; \boldsymbol\vartheta ) \right)^2}{2 \sigma^2} \right)
\end{displaymath} (2.56)

and therefore the MLE obtained by minimizing the negative log likelihood is written as
\begin{displaymath}
\hat{\boldsymbol\vartheta}_{ML} = \argmin_{\boldsymbol\vart...
...=1}^{n} \left( y_i - g(x_i ; \boldsymbol\vartheta ) \right)^2
\end{displaymath} (2.57)

that is, the conventional least-squares solution is the maximum-likelihood estimator in the presence of zero-mean additive Gaussian noise.

The $m$ partial derivatives of the log-likelihood now form a vector $m \times 1$

\begin{displaymath}
\mathbf{u}(\boldsymbol\beta) = \dfrac{\partial \ell(\boldsy...
...dots \\ \dfrac{\partial \ell} {\partial \beta_m}
\end{bmatrix}\end{displaymath} (2.58)

The vector $\mathbf{u}(\boldsymbol\beta)$ is called the score vector (or Fisher's score function) of the log-likelihood. If the log-likelihood is concave, the maximum-likelihood estimator therefore identifies the point at which
\begin{displaymath}
\mathbf{u}( \hat{ \boldsymbol\beta } ) = \mathbf{0}
\end{displaymath} (2.59)

The moments of $\mathbf{u}(\boldsymbol\beta)$ consequently satisfy important properties: as seen above, the mean of $\mathbf{u}(\boldsymbol\beta)$ evaluated at the maximum-likelihood point is zero, and the variance-covariance matrix is
\begin{displaymath}
\var \left( \mathbf{u}(\boldsymbol\beta) \right) = \E \left[...
...ta_j \partial \beta_k} \right] = \mathcal{I}(\boldsymbol\beta)
\end{displaymath} (2.60)

The matrix $\mathcal{I}$, defined as the negative Hessian, is called the expected Fisher information matrix, and its inverse is called the observed information matrix.

In machine learning, the cost function used when training many probabilistic models coincides with the negative log-likelihood or with approximations thereof.



Subsections
Paolo medici
2026-10-01