Linear Least-Squares Regression

When $f$ is a function that is linear with respect to the parameters $\boldsymbol\beta$, the problem is called linear least-squares regression (Linear Least Squares or Ordinary Least Squares, OLS). This function can be represented in the form of a linear system:

\begin{displaymath}
y_i = \mathbf{x}_i \boldsymbol\beta + \varepsilon_i
\end{displaymath} (4.13)

where $\boldsymbol\beta$ are the unknown parameters to be estimated and $\varepsilon_i$ is zero-mean additive white Gaussian noise. The parameters $\boldsymbol\beta$ are the regression coefficients: they measure the association between the variable $\mathbf{x}_i$ and the variable $y_i$. Each observation constitutes a constraint, and all individual constraints can be collected in matrix form:
\begin{displaymath}
\mathbf{y} = \mathbf{X} \boldsymbol\beta + \boldsymbol\varepsilon
\end{displaymath} (4.14)

$\mathbf{y} \in \mathbb{R}^n$ is the vector of responses (dependent variables), the matrix $\mathbf{X} \in \mathbb{R}^{n \times m}$ containing the independent variables (explanatory variables) is called the design matrix, and finally $\boldsymbol\varepsilon$ is the zero-mean additive-noise vector $\E[\boldsymbol\varepsilon]=0$ with covariance matrix $\boldsymbol\Sigma$. The parameter vector $\boldsymbol\beta$ is called the Linear Projection Coefficient or Linear Predictor. The random variable $\mathbf{y}$ therefore consists of a deterministic component and a stochastic component.

The goal is to find the hyperplane $\boldsymbol\beta$ in $m$ dimensions that best fits the data $(\mathbf{y},\mathbf{X})$.

The value $\boldsymbol\beta$ that minimizes the cost function defined in Equation (4.7), restricted to the case of observation noise with zero mean and constant variance across all samples, is in fact the best linear estimator that minimizes the variance (Best Linear Unbiased Estimator, BLUE).

Definizione 11   The Best Linear Unbiased Estimator (BLUE) of a parameter $\boldsymbol\beta$ based on a data set $Y$ is
  1. a linear estimator of the observations $Y$, so that the estimator can be written as $\hat{\boldsymbol\beta} = \mathbf{A} Y$;
  2. it must be unbiased ( $\E[\mathbf{A}Y] = \boldsymbol\beta$),
  3. among all possible linear estimators, it is the one that yields the smallest variance.

The Gauss-Markov theorem proves that a least-squares estimator is the best choice among all minimum-variance BLUE estimators when the observation variance is constant (homoscedastic).

The best least-squares estimate $\hat{\boldsymbol\beta}$ that minimizes the sum of the residuals is the solution to the linear problem

\begin{displaymath}
\hat{\boldsymbol\beta} = \argmin_\mathbf{b} \Vert \boldsymbo...
...mathbf{X}^{\top} \mathbf{X})^{-1} \mathbf{X}^{\top} \mathbf{y}
\end{displaymath} (4.15)

The same result was already obtained in Section 1.1 on the pseudoinverse of a matrix: an SVD decomposition of matrix $\mathbf{X}$ also returns the best solution in terms of computational error propagation.

The matrix $\mathbf{P}$, defined as

\begin{displaymath}
\mathbf{P} = \mathbf{X} (\mathbf{X}^{\top} \mathbf{X} )^{-1} \mathbf{X}^{\top}
\end{displaymath} (4.16)

is a projection matrix (projection matrix) that transforms the outputs (response vector) $\mathbf{y}$ into their estimates $\hat{\mathbf{y}}$ (the estimate of the noise-free observation):
\begin{displaymath}
\mathbf{P}\mathbf{y}
=
\mathbf{X}\hat{\boldsymbol\beta}
=
\hat{\mathbf{y}}.
\end{displaymath} (4.17)

By virtue of this property, $\mathbf{P}$ is called the hat matrix.

When the noise variance is not constant across the observed samples (heteroscedastic), weighted least-squares regression (Weighted Least Squares) is the BLUE choice.


\begin{displaymath}
w_i = \frac{1}{\sigma_i}
\end{displaymath} (4.18)

where $\sigma_i$ represents the standard deviation associated with observation $y_i$ and $w_i>0$ the corresponding weight.

By placing these weights in a diagonal matrix

\begin{displaymath}
\mathbf W =
\operatorname{diag}(w_1,\ldots,w_n)
\end{displaymath} (4.19)

a new linear system is obtained in which all observations have unit variance.

The solution can be expressed as

\begin{displaymath}
\hat{\boldsymbol\beta}
=
(\mathbf W \mathbf X)^+
\mathbf W \mathbf y
\end{displaymath} (4.20)

with

\begin{displaymath}
\mathbf W^\top \mathbf W
=
\boldsymbol\Sigma^{-1}.
\end{displaymath} (4.21)

More generally, in the case of noise with nonconstant variance across the observed samples and correlated samples, the best BLUE estimate in the linear case must account for the noise covariance $\boldsymbol\Sigma$:

\begin{displaymath}
\hat{\boldsymbol\beta} = (\mathbf{X}^{\top} \boldsymbol\Sigm...
...hbf{X})^{-1} \mathbf{X}^{\top}\boldsymbol\Sigma^{-1}\mathbf{y}
\end{displaymath} (4.22)

This estimator is called Generalized Least Squares (GLS).

This system minimizes the variance

\begin{displaymath}
Var[\hat{\boldsymbol\beta}_{GLS}] = (\mathbf{X}^{\top} \boldsymbol\Sigma^{-1} \mathbf{X})^{-1}
\end{displaymath} (4.23)

Paolo medici
2026-10-06