Least-Squares Regression

We first examine the most common case in real-world applications, in which the observation noise is additive white Gaussian noise.

Let

\begin{displaymath}
y = f(\mathbf{x}, \boldsymbol\beta) + \varepsilon
\end{displaymath} (4.2)

be a function, generally nonlinear, of some parameters $\boldsymbol\beta$ and some inputs $\mathbf {x}$, to which additive zero-mean Gaussian noise with variance $\sigma$ is added. To estimate the parameters robustly, the number of input samples $\mathbf{x}=\left\{ \mathbf{x}_1 \ldots \mathbf{x}_n \right\}$ must be large, much larger than the number of parameters.

The parameter function need not be the same for all samples; different functions may instead be used, effectively measuring different quantities that are all functions of the same parameters $\boldsymbol\beta$. In this case, equation (4.2) can be generalized as

\begin{displaymath}
y_i = f_i(\boldsymbol\beta) + \varepsilon_i
\end{displaymath} (4.3)

where the subscript $i$ implicitly denotes both the type of function and the $i$-th input sample (in fact, a constant parameter of the function).

The vector $\mathbf{r}$ is introduced, defined as

\begin{displaymath}
%r_i = y_i - f_i(\mathbf{x}_i, \boldsymbol\beta)
r_i = y_i - f_i(\boldsymbol\beta)
\end{displaymath} (4.4)

and containing the residual associated with the $i$-th observation (or with the $i$-th function). $r_i$ is a function of $\boldsymbol\beta$, just as $f_i$ is, and shares its derivatives with it (apart from a sign under this notation).

To obtain a maximum-likelihood estimator, the quantity to minimize is the negative log likelihood (Section 2.7) of function (4.2). For Gaussian noise, the likelihood function is in fact written as

\begin{displaymath}
\mathcal{L}(r_i \vert \boldsymbol\beta, \sigma) = \frac{1}{\sqrt{2 \pi \sigma_i^2}} e^{ - \frac{r_i^2}{2\sigma_i^2} }
\end{displaymath} (4.5)

for independent observations. Applying the definition of negative log likelihood to the likelihood function shows that, in the case of Gaussian noise, the maximum-likelihood estimator is the least-squares method.

Least-squares regression is a standard optimization technique for overdetermined systems that identifies the parameters $\boldsymbol\beta=(\beta_1, \ldots, \beta_m)$ of a function $f(\mathbf{x},\boldsymbol\beta): \mathbb{R}^{m} \mapsto \mathbb{R}^{n}$ minimizing an error $S$ computed as the sum of the squares (Sum Of Squared Error) of the residuals $r_i$ over a set of $n$ observations $y_1 \ldots y_n$:

\begin{displaymath}
S(\boldsymbol\beta) = SSE(\boldsymbol\beta) = \mathbf{r} \cd...
...= \sum_{i=1}^{n} { \Vert y_i - f_i(\boldsymbol\beta) \Vert^2 }
\end{displaymath} (4.6)

$S(\boldsymbol\beta)$ is defined as the residual sum of squares or, alternatively, as the expected squared error.

$S:\mathbb{R}^{m} \mapsto \mathbb{R}$ is a function analyzed as the parameters $\boldsymbol\beta \in \mathbb{R}^{m}$ vary in order to find its minimum value

\begin{displaymath}
\beta^{+} = \argmin_\beta S(\beta)
\end{displaymath} (4.7)

For this reason, it is called the objective function or cost function. A minimum obtained through a procedure such as the one described by equation (4.7) is called a global minimum.

From a purely computational standpoint, finding a global minimum is difficult, and techniques that identify only local minima are normally used.

Let $S(\boldsymbol\beta)$4.1 be differentiable, that is, let $f$ be differentiable. A necessary condition for $\boldsymbol\beta$ to be a minimum is that, at that point in parameter space, the gradient of $S(\boldsymbol\beta)$ vanish, namely

\begin{displaymath}
\frac{\partial S(\boldsymbol\beta)}{\partial \beta_j} = 2 \m...
...i(\boldsymbol\beta)}{\partial \beta_j} = 0 \qquad j=1,\ldots,m
\end{displaymath} (4.8)

A sufficient condition for a stationary point ( $S'(\boldsymbol\beta)=0$) to be a minimum is that $S''(\boldsymbol\beta)$ (the Hessian) be positive definite. Clearly, the existence of a local minimum only guarantees that there exists a neighborhood $\delta$ of $\boldsymbol\beta$ such that the function $S(\boldsymbol\beta + \delta) \geq S(\boldsymbol\beta)$.

The discussion so far has assumed that the noise is additive $\varepsilon$ with constant variance across all samples (homoscedasticity). When the measurement noise is still zero-mean additive Gaussian noise but has nonconstant variance, each observation $y_i$ is an independent random variable associated with variance $\sigma^{2}_i$. It is intuitive that, in this case, the optimal regression should assign greater weight to samples with low variance and lower weight to samples with high variance. To obtain this result, a normalization is used, similar to that shown in Section 2.4.1 and resulting directly from the likelihood in equation (4.5). Therefore, the simple sum of squared residuals must no longer be minimized; instead, the weighted sum of residuals must be minimized:

\begin{displaymath}
\chi^{2} = \sum^{n}_{i=1} \frac{ \Vert r_i \Vert^2 } { \sigma_i }
\end{displaymath} (4.9)

The cost function, now the square of a unit-variance random variable, becomes a chi-square distribution and is therefore denoted by $\chi^{2}$. The minimum of this cost function coincides with that previously obtained by least squares when the variance is constant. Condition (4.8) for obtaining the minimum is modified accordingly:
\begin{displaymath}
%\sum_{i=1}^{n} \frac{r_i}{\sigma_i} \frac{\partial f(\math...
...i(\boldsymbol\beta)}{\partial \beta_j} = 0 \qquad j=1,\ldots,m
\end{displaymath} (4.10)

Generalizing this concept further, when the observation is affected by Gaussian noise with known covariance matrix $\boldsymbol\Sigma$, the Weighted Sum of Squared Error (WSSE) can finally be written as

\begin{displaymath}
\chi^{2} = \mathbf{r}^{\top} \boldsymbol\Sigma^{-1} \mathbf{r}
\end{displaymath} (4.11)

It should be noted that this formulation of the cost function is equivalent to that in equation (4.6), except that the Mahalanobis distance (Section 2.4) is used instead of the Euclidean distance.

Any Weighted Least Squares problem can be reduced to an unweighted problem $\boldsymbol\Sigma = I$ by premultiplying the residuals $\mathbf{r}$ (and consequently the derivatives) by a matrix $\mathbf{L}^{\top}$ such that $\boldsymbol\Sigma^{-1} = \mathbf{L} \mathbf{L}^{\top}$, using, for example, a Cholesky decomposition when this matrix is not diagonal.

All these estimators, which account for the variance of the observation, coincide with the negative log likelihood for the variable $\mathbf{y}$ perturbed by zero-mean Gaussian noise with covariance $\boldsymbol\Sigma$.



Footnotes

...#tex2html_wrap_inline14834#4.1
In the literature, the function $S$ is often defined with a scaling factor $1/2$ so that the gradient of $S$ is not affected by the factor $2$ and has the same sign as $f$, thereby simplifying the notation.


Subsections
Paolo medici
2026-10-01