When
is a function that is linear with respect to the parameters
, the problem is called linear least-squares regression (Linear Least Squares or Ordinary Least Squares, OLS).
This function can be represented in the form of a linear system:
 |
(4.13) |
where
are the unknown parameters to be estimated and
is zero-mean additive white Gaussian noise.
The parameters
are the regression coefficients: they measure the association between the variable
and the variable
.
Each observation constitutes a constraint, and all individual constraints can be collected in matrix form:
 |
(4.14) |
is the vector of responses (dependent variables),
the matrix
containing the independent variables (explanatory variables) is called the design matrix, and finally
is the zero-mean additive-noise vector
with covariance matrix
.
The parameter vector
is called the Linear Projection Coefficient or Linear Predictor.
The random variable
therefore consists of a deterministic component and a stochastic component.
The goal is to find the hyperplane
in
dimensions that best fits the data
.
The value
that minimizes the cost function defined in Equation (4.7), restricted to the case of observation noise with zero mean and constant variance across all samples, is in fact the best linear estimator that minimizes the variance (Best Linear Unbiased Estimator, BLUE).
Definizione 11
The Best Linear Unbiased Estimator (BLUE) of a parameter
based on a data set
is
- a linear estimator of the observations
, so that the estimator can be written as
;
- it must be unbiased (
),
- among all possible linear estimators, it is the one that yields the smallest variance.
The Gauss-Markov theorem proves that a least-squares estimator is the best choice among all minimum-variance BLUE estimators when the observation variance is constant (homoscedastic).
The best least-squares estimate
that minimizes the sum of the residuals is the solution to the linear problem
 |
(4.15) |
The same result was already obtained in Section 1.1 on the pseudoinverse of a matrix: an SVD decomposition of matrix
also returns the best solution in terms of computational error propagation.
The matrix
, defined as
 |
(4.16) |
is a projection matrix (projection matrix) that transforms the outputs (response vector)
into their estimates
(the estimate of the noise-free observation):
 |
(4.17) |
By virtue of this property,
is called the hat matrix.
When the noise variance is not constant across the observed samples
(heteroscedastic), weighted least-squares regression
(Weighted Least Squares) is the BLUE choice.
 |
(4.18) |
where
represents the standard deviation associated
with observation
and
the corresponding weight.
By placing these weights in a diagonal matrix
 |
(4.19) |
a new linear system is obtained in which all observations have unit variance.
The solution can be expressed as
 |
(4.20) |
with
 |
(4.21) |
More generally, in the case of noise with nonconstant variance across the observed samples and correlated samples, the best BLUE estimate in the linear case must account for the noise covariance
:
 |
(4.22) |
This estimator is called Generalized Least Squares (GLS).
This system minimizes the variance
![\begin{displaymath}
Var[\hat{\boldsymbol\beta}_{GLS}] = (\mathbf{X}^{\top} \boldsymbol\Sigma^{-1} \mathbf{X})^{-1}
\end{displaymath}](img875.svg) |
(4.23) |
Paolo medici
2026-10-06