Linear Least-Squares Regression

When $f$ is a function that is linear with respect to the parameters $\boldsymbol\beta$, the problem is called linear least-squares regression (Linear Least Squares or Ordinary Least Squares, OLS). This function can be represented in the form of a linear system

\begin{displaymath}
y_i = \mathbf{x}_i \boldsymbol\beta + \varepsilon_i
\end{displaymath} (4.12)

where $\boldsymbol\beta$ are the unknown parameters to be determined and $\varepsilon_i$ is zero-mean additive white Gaussian noise. The parameters $\boldsymbol\beta$ are the regression coefficients: they measure the association between the variable $\mathbf {x}$ and the variable $y$. Each observation is a constraint, and all individual constraints can be collected in matrix form:
\begin{displaymath}
\mathbf{y} = \mathbf{X} \boldsymbol\beta + \boldsymbol\varepsilon
\end{displaymath} (4.13)

$\mathbf{y} \in \mathbb{R}^n$ is the vector of responses (dependent variables), the matrix $\mathbf{X} \in \mathbb{R}^{n \times m}$ collecting the independent variables (explanatory variables) is called the design matrix, and finally $\boldsymbol\varepsilon$ is the zero-mean additive-noise vector $\E[\boldsymbol\varepsilon]=0$ with variance $\boldsymbol\Sigma$. The parameter vector $\boldsymbol\beta$ is called the Linear Projection Coefficient or Linear Predictor. The random variable $\mathbf{y}$ therefore consists of a deterministic component and a stochastic component.

The goal is to find the hyperplane $\boldsymbol\beta$ in $m$ dimensions that best fits the data $(\mathbf{y},\mathbf{X})$.

The value $\boldsymbol\beta$ that minimizes the cost function defined in equation (4.6), restricted to the case of observation noise with zero mean and constant variance across all samples, is in fact the best linear estimator minimizing the variance (Best Linear Unbiased Estimator, BLUE).

Definizione 11   The Best Linear Unbiased Estimate (BLUE) of a parameter $\boldsymbol\beta$ based on a data set $Y$ is
  1. a linear function of $Y$, so that the estimator can be written as $\hat{\boldsymbol\beta} = \mathbf{A} Y$;
  2. it must be unbiased ( $\E [\mathbf{A} Y]=0$),
  3. among all possible linear estimators, it is the one that yields the smallest variance.

The Gauss-Markov theorem shows that a least-squares estimator is the best choice among all minimum-variance BLUE estimators when the observation variance is constant (homoscedastic).

The best least-squares estimate $\hat{\boldsymbol\beta}$ that minimizes the sum of the residuals is the solution of the linear problem

\begin{displaymath}
\hat{\boldsymbol\beta} = \argmin_\mathbf{b} \Vert \boldsymbo...
...mathbf{X}^{\top} \mathbf{X})^{-1} \mathbf{X}^{\top} \mathbf{y}
\end{displaymath} (4.14)

The same result was already obtained in Section 1.1 concerning the pseudoinverse of a matrix: an SVD decomposition of matrix $\mathbf{X}$ also returns the best solution in terms of computational-error propagation.

The matrix $\mathbf{P}$, defined as

\begin{displaymath}
\mathbf{P} = \mathbf{X} (\mathbf{X}^{\top} \mathbf{X} )^{-1} \mathbf{X}^{\top}
\end{displaymath} (4.15)

is a projection matrix (projection matrix) that transforms the outputs (response vector) $\mathbf{y}$ into their estimate $\hat{\mathbf{y}}$ (the estimate of the noise-free observation):
\begin{displaymath}
\mathbf{P}\mathbf{y}_i = \mathbf{x}_i \hat{\boldsymbol\beta} = \hat{\mathbf{y}}_i
\end{displaymath} (4.16)

Because of this property, $\mathbf{P}$ is called the hat matrix.

In the case of noise with nonconstant variance among the observed samples (heteroscedastic), weighted least-squares regression is the BLUE choice

\begin{displaymath}
w_i = \frac{1}{\sigma_i}
\end{displaymath} (4.17)

with $w_i > 0$ accounting for the various uncertainties associated with each observation $y_i$, such that $1/w_i$ is the standard deviation of the $i$-th measurement. After inserting the weights $w_i$ into a diagonal matrix $\mathbf{W}$, a new linear system is obtained in which each row effectively has the same observation variance. The solution that minimizes $\boldsymbol\varepsilon$ can always be expressed as
\begin{displaymath}
\hat{\boldsymbol\beta} = (\mathbf{W}\mathbf{X})^{+} \mathbf{W} \mathbf{y}
\end{displaymath} (4.18)

with $\mathbf{W}=\boldsymbol\Sigma^{-1}$.

Generalizing further, in the case of noise with nonconstant variance among the observed samples and correlations between them, the best linear BLUE estimate must account for the noise covariance $\boldsymbol\Sigma$:

\begin{displaymath}
\hat{\boldsymbol\beta} = (\mathbf{X}^{\top} \boldsymbol\Sigm...
...hbf{X})^{-1} \mathbf{X}^{\top}\boldsymbol\Sigma^{-1}\mathbf{y}
\end{displaymath} (4.19)

This estimator is called Generalized Least Squares (GLS).

This system minimizes the variance

\begin{displaymath}
Var[\hat{\boldsymbol\beta}_{GLS}] = (\mathbf{X}^{\top} \boldsymbol\Sigma^{-1} \mathbf{X})^{-1}
\end{displaymath} (4.20)

Paolo medici
2026-10-01