We first examine the most common case in real-world applications, in which the observation noise is additive white Gaussian noise.
Thus, let
The parameter function need not be the same for all samples; there may instead be different functions, each measuring a different quantity while always depending on the same parameters
.
In this case, Equation (4.3) can be generalized as
We introduce the vector defined as
| (4.5) |
To obtain a maximum-likelihood estimator, the quantity to be minimized is the negative log likelihood (Section 2.7) of function (4.3).
In the case of Gaussian noise, the likelihood function is in fact written as
Least-squares regression is a standard optimization technique for overdetermined systems that identifies the parameters
of a function
that minimize an error
computed as the sum of the squares (Sum Of Squared Errors) of the residuals
over a set of
observations
:
is a function that is analyzed as the parameters
vary in order to find its minimum value:
From a purely computational standpoint, a global minimum is difficult to identify, and techniques that identify only local minima are normally used.
Let
4.1 be differentiable, that is, let
be differentiable.
The necessary condition for
to be a minimum is that, at that point in parameter space, the gradient of
vanish, namely,
A sufficient condition for a stationary point (
) to be a minimum is that
(the Hessian) be positive definite.
Clearly, the existence of a local minimum guarantees only that there exists a neighborhood
of
such that the function
.
The discussion so far has assumed that the noise is additive with constant variance across all samples (homoscedasticity).
When the measurement noise is still zero-mean additive Gaussian noise but has nonconstant variance, each observation
is an independent random variable associated with variance
.
Intuitively, the optimal regression in this case should assign greater weight to samples with low variance and less weight to samples with high variance.
To obtain this result, a normalization is used, similar to that shown in Section 2.4.1 and directly resulting from the likelihood in Equation (4.6); consequently, one must minimize not the simple sum of squared residuals, but rather the weighted sum of residuals:
| (4.10) |
Extending this concept further, when the observation is affected by Gaussian noise with known covariance matrix
, the Weighted Sum of Squared Errors (WSSE) can finally be written as
| (4.12) |
Any Weighted Least Squares problem can be reduced to an unweighted problem
by premultiplying the residuals
(and consequently the derivatives) by a matrix
such that
, using, for example, a Cholesky decomposition when this matrix is not diagonal.
All these estimators, which account for the observation variance, coincide with the negative log likelihood for the variable perturbed by zero-mean Gaussian noise with covariance
.