The Maximum Likelihood Estimator

From a statistical standpoint, the data $\mathbf{x} = \left\{ x_1,\ldots,x_n \right\}$ can be interpreted as a sample of $n$ independent realizations of a random variable drawn from an unknown population. The task of data analysis is to identify the population that most probably generated those samples. In statistics, each population is identified by a corresponding probability distribution, and each probability distribution has a unique parameterization $\boldsymbol\vartheta$: changing these parameters must produce a different probability distribution.

Let $f( \mathbf{x} \vert \boldsymbol\vartheta)$ be the probability density function (PDF) indicating the probability of observing $\mathbf {x}$ given a parameterization $\boldsymbol\vartheta$. If the individual observations $x_i$ are statistically independent of one another, the PDF of $\mathbf {x}$ can be expressed as the product of the individual PDFs:

\begin{displaymath}
f( \mathbf{x} = \left\{ x_1 \ldots x_n \right\} \vert \bold...
...oldsymbol\vartheta) \ldots f_n(x_n \vert \boldsymbol\vartheta)
\end{displaymath} (2.51)

Given a parameterization $\boldsymbol\vartheta$, it is possible to define a specific PDF that expresses the probability of observing some data rather than other data. In the real-world case, we face the inverse problem: the data have been observed, and we must determine which $\boldsymbol\vartheta$ generated that specific PDF.

Definizione 9   To solve the inverse problem, we define the function $\mathcal{L}: \boldsymbol\vartheta \mapsto [0, \infty)$, called the likelihood function, as
\begin{displaymath}
\mathcal{L}(\boldsymbol\vartheta \vert \mathbf{x} ) = f (\m...
...rtheta) = \prod_{i=1}^{n} f_i (x_i \vert \boldsymbol\vartheta)
\end{displaymath} (2.52)

for statistically independent observations.

$\mathcal{L}( \boldsymbol\vartheta \vert \mathbf{x} )$ denotes the likelihood of the parameter $\boldsymbol\vartheta$ given the observation of the events $\mathbf {x}$.

The maximum likelihood estimator (MLE) principle $\hat{\boldsymbol\vartheta}_{MLE}$, originally developed by R. A. Fisher in the 1920s, selects as the best parameterization the one that provides the best fit between the probability distribution generated and the observed data.

For a Gaussian probability distribution, an additional definition is useful.

Definizione 10   Let $\ell$ be the log-likelihood function, defined as
\begin{displaymath}
\ell = \log \mathcal{L}(\boldsymbol\vartheta \vert x_1 \ldo...
..._n) = \sum_{i=1}^{n} \log f_i (x_i \vert \boldsymbol\vartheta)
\end{displaymath} (2.53)

where the properties of the logarithm have been used.

The best estimate of the model parameters is the one that maximizes the likelihood, or equivalently the log-likelihood,

\begin{displaymath}
\hat{\boldsymbol\vartheta}_{ML} = \argmax_{\boldsymbol\vart...
...heta} \sum_{i=1}^{n} \log f_i (x_i \vert \boldsymbol\vartheta)
\end{displaymath} (2.54)

since the logarithm is a monotonically increasing function.

In the literature, the minimum of the opposite function may be used as an optimal estimator instead of the maximum of the likelihood function:

\begin{displaymath}
\hat{\boldsymbol\vartheta}_{ML} = \argmin_{\boldsymbol\vart...
...um_{i=1}^{n} \log f_i (x_i \vert \boldsymbol\vartheta) \right)
\end{displaymath} (2.55)

that is, the minimum of the negative log-likelihood.

This formulation is particularly useful when the noise distribution is Gaussian. Let $(x_i,y_i)$ be the realizations of the random variable. For a generic function $y_i = g(x_i ; \boldsymbol\vartheta) + \epsilon$ with normally distributed noise, constant variance, and zero mean, the likelihood is

\begin{displaymath}
\mathcal{L}(\boldsymbol\vartheta \vert \mathbf{x} ) = \prod...
...- g(x_i; \boldsymbol\vartheta ) \right)^2}{2 \sigma^2} \right)
\end{displaymath} (2.56)

and therefore the MLE obtained by minimizing the negative log-likelihood is
\begin{displaymath}
\hat{\boldsymbol\vartheta}_{ML} = \argmin_{\boldsymbol\vart...
...=1}^{n} \left( y_i - g(x_i ; \boldsymbol\vartheta ) \right)^2
\end{displaymath} (2.57)

that is, the traditional least-squares solution is the maximum likelihood estimator in the presence of zero-mean additive Gaussian noise.

The $m$ partial derivatives of the log-likelihood form a vector $m \times 1$

\begin{displaymath}
\mathbf{u}(\boldsymbol\beta) = \dfrac{\partial \ell(\boldsy...
...dots \\ \dfrac{\partial \ell} {\partial \beta_m}
\end{bmatrix}\end{displaymath} (2.58)

The vector $\mathbf{u}(\boldsymbol\beta)$ is called the score vector (or Fisher's score function) of the log-likelihood. If the log-likelihood is concave, the maximum likelihood estimator therefore identifies the point at which
\begin{displaymath}
\mathbf{u}( \hat{ \boldsymbol\beta } ) = \mathbf{0}
\end{displaymath} (2.59)

The moments of $\mathbf{u}(\boldsymbol\beta)$ therefore satisfy important properties: as seen above, the mean of $\mathbf{u}(\boldsymbol\beta)$ evaluated at the maximum-likelihood point is zero, and the covariance matrix is
\begin{displaymath}
\var \left( \mathbf{u}(\boldsymbol\beta) \right) = \E \left[...
...ta_j \partial \beta_k} \right] = \mathcal{I}(\boldsymbol\beta)
\end{displaymath} (2.60)

The matrix $\mathcal{I}$, defined as the expected value of the negative Hessian of the log-likelihood, is called the expected Fisher information matrix. The negative Hessian of the log-likelihood evaluated on the observed data is instead called the observed information matrix.

In machine learning, the cost function used during the training of many probabilistic models coincides with the negative log-likelihood or with approximations thereof.



Subsections
Paolo medici
2026-10-06