From a statistical standpoint, the data vector
consists of realizations of a random variable from an unknown population.
The task of data analysis is to identify the population most likely to have generated these samples. In statistics, each population is identified by a corresponding probability distribution, and each probability distribution is associated with a unique parameterization
: varying these parameters must produce a different probability distribution.
Let
be the probability density function (PDF) indicating the probability of observing
given a parameterization
.
If the individual observations
are statistically independent of one another, the PDF of
can be expressed as the product of the individual PDFs:
| (2.51) |
Given a parameterization
, it is possible to define a specific PDF that expresses the probability of some data occurring rather than other data.
In the real-world case, we face the inverse problem: the data have been observed, and we must determine which
generated that specific PDF.
| (2.52) |
The maximum-likelihood estimator (MLE) principle
, originally developed by R. A. Fisher in the 1920s, selects as the best parameterization the one that provides the best fit of the probability distribution generated by the observed data.
For a Gaussian probability distribution, an additional definition is useful.
| (2.53) |
The best estimate of the model parameters is the one that maximizes the likelihood, or equivalently the log-likelihood
| (2.54) |
In the literature, the optimum estimator is sometimes defined not as the maximum of the likelihood function, but as the minimum of its negative
| (2.55) |
This formulation is particularly useful when the noise distribution is Gaussian.
Let be realizations of the random variable.
In fact, for a generic function
with normally distributed noise, constant variance, and zero mean, the likelihood is
| (2.56) |
| (2.57) |
The partial derivatives of the log-likelihood now form a vector
| (2.58) |
| (2.59) |
| (2.60) |
In machine learning, the cost function used when training many probabilistic models coincides with the negative log-likelihood or with approximations thereof.