Least-Squares Regression

We first examine the most common case in real-world applications, in which the observation noise is additive white Gaussian noise.

Thus, let

\begin{displaymath}
y = f(\mathbf{x}, \boldsymbol\beta) + \varepsilon
\end{displaymath} (4.3)

be a function, generally nonlinear, of some parameters $\boldsymbol\beta$ and some inputs $\mathbf {x}$, to which additive zero-mean Gaussian noise with variance $\sigma^2$ is added. To estimate the parameters robustly, the number of input samples $\mathbf{x}=\left\{ \mathbf{x}_1 \ldots \mathbf{x}_n \right\}$ must be large, much larger than the number of parameters.

The parameter function need not be the same for all samples; there may instead be different functions, each measuring a different quantity while always depending on the same parameters $\boldsymbol\beta$. In this case, Equation (4.3) can be generalized as

\begin{displaymath}
y_i = f_i(\boldsymbol\beta) + \varepsilon_i
\end{displaymath} (4.4)

where the subscript $i$ implicitly denotes both the function type and the \(i\)-th input sample (in effect, a constant parameter of the function).

We introduce the vector $\mathbf{r}$ defined as

\begin{displaymath}
%r_i = y_i - f_i(\mathbf{x}_i, \boldsymbol\beta)
r_i = y_i - f_i(\boldsymbol\beta)
\end{displaymath} (4.5)

containing the residual associated with the \(i\)-th observation (or the \(i\)-th function). $r_i$ is a function of $\boldsymbol\beta$ just as $f_i$ is, and shares its derivatives with it (up to a sign under this convention).

To obtain a maximum-likelihood estimator, the quantity to be minimized is the negative log likelihood (Section 2.7) of function (4.3). In the case of Gaussian noise, the likelihood function is in fact written as

\begin{displaymath}
\mathcal{L}(r_i \vert \boldsymbol\beta, \sigma) = \frac{1}{\sqrt{2 \pi \sigma_i^2}} e^{ - \frac{r_i^2}{2\sigma_i^2} }
\end{displaymath} (4.6)

for independent observations. Applying the definition of negative log likelihood to the likelihood function shows that, in the case of Gaussian noise, the maximum-likelihood estimator is the least-squares method.

Least-squares regression is a standard optimization technique for overdetermined systems that identifies the parameters $\boldsymbol\beta=(\beta_1, \ldots, \beta_m)$ of a function $f(\mathbf{x},\boldsymbol\beta): \mathbb{R}^{m} \mapsto \mathbb{R}^{n}$ that minimize an error $S$ computed as the sum of the squares (Sum Of Squared Errors) of the residuals $r_i$ over a set of $n$ observations $y_1 \ldots y_n$:

\begin{displaymath}
S(\boldsymbol\beta) = SSE(\boldsymbol\beta) = \mathbf{r} \cd...
...= \sum_{i=1}^{n} { \Vert y_i - f_i(\boldsymbol\beta) \Vert^2 }
\end{displaymath} (4.7)

$S(\boldsymbol\beta)$ is defined as the residual sum of squares, or alternatively as the expected squared error.

$S:\mathbb{R}^{m} \mapsto \mathbb{R}$ is a function that is analyzed as the parameters $\boldsymbol\beta \in \mathbb{R}^{m}$ vary in order to find its minimum value:

\begin{displaymath}
\beta^{+} = \argmin_\beta S(\beta)
\end{displaymath} (4.8)

For this reason, it is called the objective function or cost function. A minimum obtained through a procedure such as that described by Equation (4.8) is called a global minimum.

From a purely computational standpoint, a global minimum is difficult to identify, and techniques that identify only local minima are normally used.

Let $S(\boldsymbol\beta)$4.1 be differentiable, that is, let $f$ be differentiable. The necessary condition for $\boldsymbol\beta$ to be a minimum is that, at that point in parameter space, the gradient of $S(\boldsymbol\beta)$ vanish, namely,

\begin{displaymath}
\frac{\partial S(\boldsymbol\beta)}{\partial \beta_j} = 2 \m...
...i(\boldsymbol\beta)}{\partial \beta_j} = 0 \qquad j=1,\ldots,m
\end{displaymath} (4.9)

A sufficient condition for a stationary point ( $S'(\boldsymbol\beta)=0$) to be a minimum is that $S''(\boldsymbol\beta)$ (the Hessian) be positive definite. Clearly, the existence of a local minimum guarantees only that there exists a neighborhood $\delta$ of $\boldsymbol\beta$ such that the function $S(\boldsymbol\beta + \delta) \geq S(\boldsymbol\beta)$.

The discussion so far has assumed that the noise is additive $\varepsilon$ with constant variance across all samples (homoscedasticity). When the measurement noise is still zero-mean additive Gaussian noise but has nonconstant variance, each observation $y_i$ is an independent random variable associated with variance $\sigma^{2}_i$. Intuitively, the optimal regression in this case should assign greater weight to samples with low variance and less weight to samples with high variance. To obtain this result, a normalization is used, similar to that shown in Section 2.4.1 and directly resulting from the likelihood in Equation (4.6); consequently, one must minimize not the simple sum of squared residuals, but rather the weighted sum of residuals:

\begin{displaymath}
\chi^{2} = \sum^{n}_{i=1} \frac{ \Vert r_i \Vert^2 } { \sigma_i^2 }
\end{displaymath} (4.10)

The cost function, obtained as the sum of the squares of standardized residuals,4.2 is commonly denoted by the symbol $\chi^{2}$. The minimum of this cost function coincides with that previously obtained by least squares when the variance is constant. Condition (4.9) for obtaining the minimum is modified accordingly:
\begin{displaymath}
\sum_{i=1}^{n} \frac{r_i}{\sigma_i^2} \frac{\partial f_i(\boldsymbol\beta)} {\partial \beta_j} = 0 \qquad j=1,\ldots,m
\end{displaymath} (4.11)

Extending this concept further, when the observation is affected by Gaussian noise with known covariance matrix $\boldsymbol\Sigma$, the Weighted Sum of Squared Errors (WSSE) can finally be written as

\begin{displaymath}
\chi^{2} = \mathbf{r}^{\top} \boldsymbol\Sigma^{-1} \mathbf{r}
\end{displaymath} (4.12)

It should be noted that this formulation of the cost function is equivalent to that in Equation (4.7), except that the Mahalanobis distance (Section 2.4) is used instead of the Euclidean distance.

Any Weighted Least Squares problem can be reduced to an unweighted problem $\boldsymbol\Sigma = I$ by premultiplying the residuals $\mathbf{r}$ (and consequently the derivatives) by a matrix $\mathbf{L}^{\top}$ such that $\boldsymbol\Sigma^{-1} = \mathbf{L} \mathbf{L}^{\top}$, using, for example, a Cholesky decomposition when this matrix is not diagonal.

All these estimators, which account for the observation variance, coincide with the negative log likelihood for the variable $\mathbf{y}$ perturbed by zero-mean Gaussian noise with covariance $\boldsymbol\Sigma$.



Footnotes

...#tex2html_wrap_inline14850#4.1
In the literature, the function $S$ is often defined with a scale factor $1/2$ so that the gradient of $S$ is not affected by the factor $2$ and has the same sign as $f$, thereby simplifying the notation.
... residuals,4.2
Under the usual assumption of independent Gaussian noise, the resulting statistic follows a distribution $\chi^2$. When the model parameters are estimated from the data, the distribution's degrees of freedom are reduced to $n-m$.


Subsections
Paolo medici
2026-10-06