From a statistical standpoint, the data
can be interpreted as a sample of
independent realizations of a random variable drawn from an unknown population.
The task of data analysis is to identify the population that most probably generated those samples. In statistics, each population is identified by a corresponding probability distribution, and each probability distribution has a unique parameterization
: changing these parameters must produce a different probability distribution.
Let
be the probability density function (PDF) indicating the probability of observing
given a parameterization
.
If the individual observations
are statistically independent of one another, the PDF of
can be expressed as the product of the individual PDFs:
| (2.51) |
Given a parameterization
, it is possible to define a specific PDF that expresses the probability of observing some data rather than other data.
In the real-world case, we face the inverse problem: the data have been observed, and we must determine which
generated that specific PDF.
| (2.52) |
The maximum likelihood estimator (MLE) principle
, originally developed by R. A. Fisher in the 1920s, selects as the best parameterization the one that provides the best fit between the probability distribution generated and the observed data.
For a Gaussian probability distribution, an additional definition is useful.
| (2.53) |
The best estimate of the model parameters is the one that maximizes the likelihood, or equivalently the log-likelihood,
| (2.54) |
In the literature, the minimum of the opposite function may be used as an optimal estimator instead of the maximum of the likelihood function:
| (2.55) |
This formulation is particularly useful when the noise distribution is Gaussian.
Let be the realizations of the random variable.
For a generic function
with normally distributed noise, constant variance, and zero mean, the likelihood is
| (2.56) |
| (2.57) |
The partial derivatives of the log-likelihood form a vector
| (2.58) |
| (2.59) |
| (2.60) |
In machine learning, the cost function used during the training of many probabilistic models coincides with the negative log-likelihood or with approximations thereof.