When
is a function that is linear with respect to the parameters
, the problem is called linear least-squares regression (Linear Least Squares or Ordinary Least Squares, OLS).
This function can be represented in the form of a linear system
 |
(4.12) |
where
are the unknown parameters to be determined and
is zero-mean additive white Gaussian noise.
The parameters
are the regression coefficients: they measure the association between the variable
and the variable
.
Each observation is a constraint, and all individual constraints can be collected in matrix form:
 |
(4.13) |
is the vector of responses (dependent variables),
the matrix
collecting the independent variables (explanatory variables) is called the design matrix, and finally
is the zero-mean additive-noise vector
with variance
.
The parameter vector
is called the Linear Projection Coefficient or Linear Predictor.
The random variable
therefore consists of a deterministic component and a stochastic component.
The goal is to find the hyperplane
in
dimensions that best fits the data
.
The value
that minimizes the cost function defined in equation (4.6), restricted to the case of observation noise with zero mean and constant variance across all samples, is in fact the best linear estimator minimizing the variance (Best Linear Unbiased Estimator, BLUE).
Definizione 11
The Best Linear Unbiased Estimate (BLUE) of a parameter
based on a data set
is
- a linear function of
, so that the estimator can be written as
;
- it must be unbiased (
),
- among all possible linear estimators, it is the one that yields the smallest variance.
The Gauss-Markov theorem shows that a least-squares estimator is the best choice among all minimum-variance BLUE estimators when the observation variance is constant (homoscedastic).
The best least-squares estimate
that minimizes the sum of the residuals is the solution of the linear problem
 |
(4.14) |
The same result was already obtained in Section 1.1 concerning the pseudoinverse of a matrix: an SVD decomposition of matrix
also returns the best solution in terms of computational-error propagation.
The matrix
, defined as
 |
(4.15) |
is a projection matrix (projection matrix) that transforms the outputs (response vector)
into their estimate
(the estimate of the noise-free observation):
 |
(4.16) |
Because of this property,
is called the hat matrix.
In the case of noise with nonconstant variance among the observed samples (heteroscedastic), weighted least-squares regression is the BLUE choice
 |
(4.17) |
with
accounting for the various uncertainties associated with each observation
, such that
is the standard deviation of the
-th measurement.
After inserting the weights
into a diagonal matrix
, a new linear system is obtained in which each row effectively has the same observation variance.
The solution that minimizes
can always be expressed as
 |
(4.18) |
with
.
Generalizing further, in the case of noise with nonconstant variance among the observed samples and correlations between them, the best linear BLUE estimate must account for the noise covariance
:
 |
(4.19) |
This estimator is called Generalized Least Squares (GLS).
This system minimizes the variance
![\begin{displaymath}
Var[\hat{\boldsymbol\beta}_{GLS}] = (\mathbf{X}^{\top} \boldsymbol\Sigma^{-1} \mathbf{X})^{-1}
\end{displaymath}](img869.svg) |
(4.20) |
Paolo medici
2026-10-01